---
title: "Introduction to sidrar"
author: "Renato Prado Siqueira"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Introduction to sidrar}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(echo = TRUE, collapse = TRUE, comment = "#>")
```

## Overview

`sidrar` is an R interface to SIDRA (*Sistema IBGE de Recuperação
Automática*), the system through which the Brazilian Institute of Geography
and Statistics (IBGE) publishes aggregate statistics.

The usual workflow is:

1. find a table with `search_sidra()`;
2. inspect its available dimensions with `info_sidra()`; and
3. optionally build and inspect a request with `sidra_query()` and
   `sidra_plan()`;
4. retrieve a selection with `get_sidra()`; or
5. split and collect a large explicit request with `sidra_split()` and
   `sidra_collect()`.

Network-dependent examples are not evaluated while the vignette is built.

## Installation

Install the released version from CRAN:

```{r, eval = FALSE}
install.packages("sidrar")
```

Install the development version from GitHub with `pak`:

```{r, eval = FALSE}
# install.packages("pak")
pak::pak("rpradosiqueira/sidrar")
```

## Find and inspect a table

`search_sidra()` searches titles in IBGE's official aggregate catalog. Its
result is a character vector whose names are the SIDRA table codes:

```{r, eval = FALSE}
library(sidrar)

search_sidra("IPCA")
search_sidra(c("contas", "nacionais"))
```

The search is case- and accent-insensitive. When several terms are supplied,
all terms must occur in the title, but they need not be adjacent.

Once you have a code, inspect the periods, variables, classifications,
categories, and territorial levels accepted by the table:

```{r, eval = FALSE}
metadata <- info_sidra(7060)
names(metadata)
metadata$variable
metadata$classific_category
metadata$geo
```

Set `wb = TRUE` to open the official table descriptor in the default browser.
The function no longer prompts for confirmation:

```{r, eval = FALSE}
info_sidra(7060, wb = TRUE)
```

For programmatic work, use the normalized discovery layer. It preserves codes
as character strings and returns stable base R data frames:

```{r, eval = FALSE}
catalog <- sidra_catalog()
metadata <- sidra_metadata(7060)
periods <- sidra_periods(7060)
locations <- sidra_locations(7060, "N1")

names(metadata)
metadata$variables
metadata$classifications
metadata$categories
```

The legacy `search_sidra()` and `info_sidra()` contracts remain unchanged.
If SIDRA's descriptor returns a recognized browser challenge, `info_sidra()`
uses official aggregate metadata and periods. It discloses unavailable
descriptor-specific names, geographic counts, and variable availability
exceptions in `attr(metadata, "sidrar_metadata")` when `metadata` is its
return value; it does not invent equivalent fields. `wb = TRUE` still opens
the original descriptor. `options(sidrar.fallback = FALSE)` disables this route.

## Build a structured request

This request retrieves the monthly IPCA for the general index in Campo
Grande, Mato Grosso do Sul, over the 12 most recent periods:

```{r, eval = FALSE}
ipca <- get_sidra(
  x = 7060,
  variable = 63,
  period = c(last = 12),
  geo = "City",
  geo.filter = list(City = 5002704),
  classific = "c315",
  category = list(7169)
)
```

`geo.filter` may also select every unit inside a higher territorial level.
For example, the following pattern requests cities inside Mato Grosso do Sul:

```{r, eval = FALSE}
get_sidra(
  x = 7060,
  variable = 63,
  period = "last",
  geo = "City",
  geo.filter = list(State = 50),
  classific = "c315",
  category = list(7169)
)
```

The existing defaults remain unchanged: descriptive headers are enabled,
`format = 4` requests codes and names, `digits = "default"` uses the table's
standard precision, and `variable = "allxp"` excludes automatically generated
percentage variables.

Build the same request without downloading values and inspect the selections
whose cardinality can be determined offline:

```{r, eval = FALSE}
query <- sidra_query(
  x = 7060,
  variable = 63,
  period = sprintf("2024%02d", 1:12),
  geo = "City",
  geo.filter = list(City = 5002704),
  classific = "c315",
  category = list(7169)
)

query$url
sidra_plan(query)
```

No service limit is assumed by the planner. If a limit is known for the
current request, pass it explicitly with `sidra_plan(query, limit = ...)`.

## Split and collect batches

When one dimension contains many explicit members, split it into disjoint
batches and collect them sequentially:

```{r, eval = FALSE}
batches <- sidra_split(query, by = "period", size = 6)
data <- sidra_collect(batches, provenance = TRUE)
sidra_provenance(data)
```

`sidra_collect()` requires identical names and column types across batches. It
does not sort or deduplicate rows. Category members containing a space are
SIDRA sums and are kept indivisible in structured queries. URLs can now be
split directly; period selectors `all`, `first`, `last`, and ranges use the
official inventory without inventing calendar periods. Other special
selectors require explicit codes. A geographic filter can
be split only for a query with one non-Brazil territorial level; multiple
levels would repeat the unchanged levels in every batch.

Version 0.6.0 adds opt-in period batching and resumable local checkpoints:

```{r, eval = FALSE}
url <- "/t/6468/n1/all/n2/all/n3/all/v/4099/p/all/h/n"
batches <- sidra_split(url, "period", size = 8)
batches$resolution$selection

data <- sidra_collect(
  url, batch_size = 8, value_type = "both",
  checkpoint = "sidrar-pnad", provenance = TRUE
)
# Repeat the same call after an interruption to reuse completed batches.
sidra_provenance(data)$batch_accessed_at
sidra_provenance(data)$resumed
```

`batch_size` limits periods, not cells: one period can still exceed the API
limit and require splitting another explicit dimension. `get_sidra()` remains
unchanged and does not split or store values. Checkpoints freeze the period
inventory, including relative selectors when `batch_size` is omitted, and
validate settings, package version, checksums, and schemas before reuse.
They are trusted local files, separate from the metadata cache. An existing
checkpoint is never silently cleared; use a new directory for a fresh extract.
Stored and new batches may span source revisions, so inspect access times.
The result is still combined in memory, not queried from disk. A stale
`.sidrar-lock` after a hard crash must only be removed after confirming there
is no running collector. URL geographic splitting requires direct codes;
use a structured query for containment filters.

For comparisons between services or collections, align by territorial-level,
location, period, variable, and classification codes, not row positions.
Names can change and two territorial levels can reuse a location code.

Territorial views (`G`) and extinct territorial units (`/u/y`) are available
through additive arguments:

```{r, eval = FALSE}
sidra_query(1612, geo_view = 44, classific = character())
sidra_query(
  1612,
  geo = "State",
  geo.filter = list(c(20, 34)),
  include_extinct = TRUE,
  classific = character()
)
```

## Use an API path or full URL

If a query was assembled elsewhere, pass either its path:

```{r, eval = FALSE}
get_sidra(
  api = "/t/7060/n1/all/v/63/p/last/c315/7169"
)
```

or the complete official HTTPS URL:

```{r, eval = FALSE}
get_sidra(
  api = paste0(
    "https://apisidra.ibge.gov.br/values/",
    "t/7060/n1/all/v/63/p/last/c315/7169"
  )
)
```

Full URLs are restricted to the official
`https://apisidra.ibge.gov.br/values` endpoint. Percent-encoded segments such
as `%20` are preserved. If the path contains `/h/n`, the first observation is
kept as data instead of being interpreted as a header.

## Preserve SIDRA's special values

SIDRA uses symbols with specific meanings, including `"-"` for an absolute
zero, `"X"` for an inhibited value, `".."` when a value does not apply, and
`"..."` when it is unavailable. Earlier versions returned a numeric `Valor`
column, so special symbols became `NA`. That remains the default for
compatibility.

Use `value_type = "character"` to keep the symbols directly:

```{r, eval = FALSE}
raw <- get_sidra(
  api = "/t/1849/n3/all/v/811/p/2018/c12762/all",
  value_type = "character"
)
```

Use `value_type = "both"` to keep numeric `Valor` and append `Valor_raw`:

```{r, eval = FALSE}
both <- get_sidra(
  api = "/t/1849/n3/all/v/811/p/2018/c12762/all",
  value_type = "both"
)
```

## Network behavior

Requests use HTTPS, UTF-8 decoding, an identifying user agent, a timeout, and
limited retries for transient failures. Customize the timeout and retry count
with:

```{r, eval = FALSE}
options(
  sidrar.timeout = 120,
  sidrar.retries = 4
)
```

Catalog and metadata caching is explicit and disabled by default. Value
responses are not cached automatically; explicit collection checkpoints are
separate from the metadata cache:

```{r, eval = FALSE}
metadata <- sidra_metadata(7060, cache = TRUE)
sidra_cache_info()
sidra_cache_clear()
```

Use `refresh = TRUE` to bypass and replace a cached discovery entry.

Cloudflare browser challenges are identified as `sidrar_challenge_error`.
For compatible values queries, `get_sidra()` and `sidra_collect()` can use
IBGE's official aggregate API v3 as an alternative endpoint. This fallback
supports multiple geographic levels, explicit periods and ranges, complete
or first/latest period selections, standard variable/category selections,
and the default descriptor format. Dimension columns follow the original URL;
observation order remains that returned by the alternative service. It
preserves the requested headers and special-value handling, and announces
when the alternative is used. Set `options(sidrar.fallback = FALSE)` to
disable it. Unsupported selections retain the original challenge error with
an explanation in `fallback_reason`.

Explicit decimal precision is accepted only when numeric values already have
the requested decimal places. Otherwise `sidrar_fallback_precision_error`
(also a `sidrar_parse_error`) prevents silent re-rounding or invented digits.
Default precision preserves the received values; maximum precision is not
supported by this alternative. Automatic classification discovery can use
official aggregate metadata when the SIDRA descriptor returns a challenge.

Every fallback response is checked for complete dimension fields, textual
identifiers, duplicate observation keys, and codes outside explicit filters.
Missing explicitly requested members generate `sidrar_incomplete_warning`,
not fabricated rows or zeros: legitimate sparse tables need not form a full
Cartesian product. Full coverage of `all` and contextual geographic membership
still require comparison with current metadata.

For HTTP 429 or 503, the client honors a valid `Retry-After` header, including
HTTP dates. The default maximum accepted server delay is 60 seconds; set
`options(sidrar.retry_after_max = 120)` to allow up to two minutes. The value
must be one finite positive number of seconds; invalid settings use 60, and
`options(sidrar.retry_after_max = NULL)` restores that default. If another
attempt would exceed this limit, `sidrar_retry_after_error` carries the delay
in `retry_after` and the configured limit in `retry_after_max`; it does not
trigger an early retry or change the per-attempt timeout.

```{r, eval = FALSE}
pnad <- get_sidra(
  api = "/t/6468/n1/all/n2/all/n3/all/v/4099/p/all/d/v4099%201",
  value_type = "both"
)
```

Collection provenance records actual source endpoints in `urls` and the
original queries in `requested_urls` when fallback occurs. If both official
endpoints fail, report the URL, time, and the error's `cf_ray` to IBGE;
increasing retries cannot solve an interactive browser challenge.

Invalid parameters and API limits are reported with the response returned by
SIDRA. See the
[official API help](https://apisidra.ibge.gov.br/home/ajuda) for the complete
query syntax and current service limits.
