OPCC links Ontario postal codes to Statistics Canada 2021 census geographies through open, redistributable evidence.
Two access paths:
| You want to⦠| Go to |
|---|---|
| Look up postal codes or join census geographies to your own data | Part 1 β uses pre-built artifacts downloaded automatically by the package; no source checkout needed |
| Verify artifacts or reproduce the DA roll-up from the DB artifact | Part 2 β uses package functions only; no source checkout needed |
| Rebuild all artifacts from raw Statistics Canada sources | Rebuild all
artifacts β uses package functions for download and build; requires
sf, dplyr, readr |
Most code below is shown but not run when this vignette is built.
Those chunks download versioned release artifacts, read files from paths
you supply, or need packages listed under Suggests. Chunks
that run with the installed package alone and no network are
evaluated.
OPCC links postal codes to two Statistics Canada census geographies:
DBUID.DAUID. Every DB belongs to
exactly one DA.Each postal code can map to multiple geographies. OPCC retains all
defensible links with an allocation_weight (the share of
observed address evidence for that link). The best_link
flag marks the single strongest link per postal code. Unmatched postal
codes are reported, never fabricated.
All OPCC functions normalize internally, but you can normalize explicitly:
The first call downloads and caches the artifact (~15 MB). Subsequent calls use the local cache.
# Best dissemination area (DA) link
pc_to_geo("M5V 3A8", level = "DA", all_links = FALSE)
# Every defensible DA link with allocation weights
pc_to_geo("M5V 3A8", level = "DA")
# Best dissemination block (DB) link
pc_to_geo("M5V 3A8", level = "DB", all_links = FALSE)
# Every DB link
pc_to_geo("M5V 3A8", level = "DB")Unmatched postal codes are not silently dropped. They appear in the
unmatched attribute of the result:
The most common workflow: you have a data frame with a postal-code
column and you want census geography identifiers.
get_correspondence() and
get_da_correspondence() download a release artifact, so the
joins below are not evaluated here.
library(OPCC)
my_data <- data.frame(
id = 1:5,
postal_code = c("M5V 3A8", "K1A 0A6", "N6A 1B1", "P7A 1A1", "ZZZ 9Z9"),
value = c(10, 20, 30, 40, 50),
stringsAsFactors = FALSE
)
# --- DA join (one best DA per postal code) ---
da <- get_da_correspondence(vintage = "2026-07-20")
da_best <- da[da$best_link, c("postal_code", "DAUID", "allocation_weight")]
merged_da <- merge(
my_data,
da_best,
by = "postal_code",
all.x = TRUE
)
# Rows with unmatched postal codes keep NA for DAUID and allocation_weight.
# --- DB join (one best DB per postal code) ---
db <- get_correspondence(vintage = "2026-07-19-geonames-amendment")
db_best <- db[db$best_link, c("postal_code", "DBUID", "DAUID", "allocation_weight")]
merged_db <- merge(
my_data,
db_best,
by = "postal_code",
all.x = TRUE
)
# --- Keep all many-to-many links ---
# Skip the best_link filter to retain every candidate geography and its
# allocation weight. Your join will produce multiple rows per postal code.
merged_all <- merge(my_data, da, by = "postal_code", all.x = TRUE)Columns returned by the join:
| Column | Meaning |
|---|---|
postal_code |
Normalized A1A 1A1 postal code |
DBUID |
2021 Dissemination Block identifier (DB artifact) |
DAUID |
2021 Dissemination Area identifier (both artifacts) |
allocation_weight |
Share of observed address evidence for this link; sums to 1 per postal code |
best_link |
TRUE for the single strongest link per postal code |
n_contributing_dbs |
Number of DBs contributing to a DA link (DA artifact) |
contributing_dbuids |
Pipe-delimited DBUID list behind a DA link (DA artifact) |
source_vintages |
Source evidence vintage(s) behind the link |
With dplyr, which is a Suggests package and
so is also not evaluated here:
get_correspondence() and
get_da_correspondence() download the versioned artifact,
verify its SHA-256 checksum, cache it locally, and return a data frame.
Subsequent calls use the cache.
By default, artifacts are cached in a session temporary directory, so
OPCC never writes to your home filespace unless you ask it to. To keep
artifacts across sessions, pass an explicit cache_dir, set
the OPCC.cache_dir option or the
OPCC_CACHE_DIR environment variable, or accept the one-time
prompt that an interactive session offers. See
?clear_opcc_cache for the full resolution order.
Listing vintages reads an index shipped with the package, so it runs here:
list_vintages(level = "DB")
#> [1] "2026-06-26" "2026-07-19-geonames-amendment"
list_vintages(level = "DA")
#> [1] "2026-06-26" "2026-07-20"Downloading a release does need the network, so the rest is shown but not run:
# Download using the default session cache
db <- get_correspondence(vintage = "2026-07-19-geonames-amendment")
da <- get_da_correspondence(vintage = "2026-07-20")
# Or use an explicit cache directory
cache <- file.path(tempdir(), "opcc-cache")
db <- get_correspondence(
vintage = "2026-07-19-geonames-amendment",
cache_dir = cache
)
# Pass a pre-loaded artifact to skip re-reading
pc_to_geo("M5V 3A8", level = "DA", correspondence = da)
pc_to_geo("M5V 3A8", level = "DB", correspondence = db)# Download, checksum, and validate schema and invariants
validate_release(vintage = "2026-07-19-geonames-amendment", level = "DB",
cache_dir = cache)
validate_release(vintage = "2026-07-20", level = "DA", cache_dir = cache)
# Inspect provenance metadata
release_manifest(vintage = "2026-07-20", level = "DA", cache_dir = cache)After caching a verified release once, use
offline = TRUE to require the local cache and block network
access:
validate_release(vintage = "2026-07-20", level = "DA",
cache_dir = cache, offline = TRUE)
da <- get_da_correspondence(vintage = "2026-07-20",
cache_dir = cache, offline = TRUE)You can also supply a local artifact file directly to
pc_to_point():
pc_to_point() returns supplementary GeoNames point
observations. These are source-labeled and separate from NAR address
evidence. It downloads its own point artifact, so the chunk is not
evaluated:
Use this only for redistributable, open data. Local layers stay separate from canonical OPCC releases. The API rejects Canada Post, PCCF, and PCCF+ inputs. The paths below are placeholders for your own files, so the chunk is not evaluated.
Describing the source is pure metadata, so this part runs. A real
checksum comes from your own file; a literal stands in for
it here:
adapter <- new_source_adapter(
source_id = "municipal_registry",
licence = "Open Government Licence - Municipality",
lineage = "Municipal open address registry",
retrieval_date = "2026-07-20",
schema_map = list(
postal_code = "postal",
latitude = "lat",
longitude = "lon"
),
checksum = strrep("0", 64L)
)
#> This source layer remains local and separate from canonical OPCC releases. If redistribution is permitted, submit its contribution bundle as an OPCC issue or pull request.
adapter
#> $source_id
#> [1] "municipal_registry"
#>
#> $licence
#> [1] "Open Government Licence - Municipality"
#>
#> $lineage
#> [1] "Municipal open address registry"
#>
#> $retrieval_date
#> [1] "2026-07-20"
#>
#> $schema_map
#> $schema_map$postal_code
#> [1] "postal"
#>
#> $schema_map$latitude
#> [1] "lat"
#>
#> $schema_map$longitude
#> [1] "lon"
#>
#>
#> $endpoint
#> NULL
#>
#> $checksum
#> [1] "0000000000000000000000000000000000000000000000000000000000000000"
#>
#> $location_type
#> [1] "unknown"
#>
#> $coordinate_method
#> [1] "unknown"
#>
#> $authority_level
#> [1] "unknown"
#>
#> $coverage_type
#> [1] "unknown"
#>
#> $update_frequency
#> [1] "unknown"
#>
#> attr(,"class")
#> [1] "opcc_source_adapter"Building the layer, profiling it, and packaging a contribution bundle all need your own data, so they are shown but not run:
layer <- build_source_layer(my_source, adapter, on_invalid = "quarantine")
profile_source_layer(layer)
# `output_dir` must be an explicit path you choose; nothing is written
# to the working directory by default.
bundle <- contribution_bundle(layer,
output_dir = file.path(tempdir(), "contributions"),
fixture_rows = 100L)
contribution_issue_url(bundle)Every published artifact can be downloaded, verified, and reproduced using package functions alone. No source checkout is needed.
Download each artifact, check its SHA-256 against the release index, and validate schema and correspondence invariants (unique keys, weight sums, one best link per postal code):
validate_release(vintage = "2026-07-19-geonames-amendment", level = "DB")
validate_release(vintage = "2026-07-20", level = "DA")Inspect the full provenance metadata (source checksums, code version, build timestamp, row counts):
The DA correspondence (M5) is a deterministic attribute roll-up of
the DB correspondence (M2). aggregate_da_correspondence()
performs the same computation as the M5 build script: it sums DB
allocation weights within each DA, preserves contributing-DB lineage,
and selects the best DA link.
The published DA artifacts were built from the NAR-only M2 baseline
(2026-06-26), so use that vintage to reproduce them
exactly. Both lookups download a release artifact, so the chunk is not
evaluated:
db <- get_correspondence(vintage = "2026-06-26")
da_reproduced <- aggregate_da_correspondence(db)
# Compare with the published DA artifact
da_published <- get_da_correspondence(vintage = "2026-07-20")
stopifnot(identical(nrow(da_reproduced), nrow(da_published)))
stopifnot(identical(
sort(unique(da_reproduced$postal_code)),
sort(unique(da_published$postal_code))
))The GeoNames amendment vintage
(2026-07-19-geonames-amendment) adds 17,334 supplementary
postal codes. Rolling it up with
aggregate_da_correspondence() produces a valid superset of
the published DA, but it will not match the published artifact
row-for-row.
After caching artifacts once, all verification and reproduction works without network access:
cache <- file.path(tempdir(), "opcc-cache")
validate_release(vintage = "2026-06-26", level = "DB",
cache_dir = cache, offline = TRUE)
validate_release(vintage = "2026-07-20", level = "DA",
cache_dir = cache, offline = TRUE)
db <- get_correspondence(vintage = "2026-06-26",
cache_dir = cache, offline = TRUE)
da <- aggregate_da_correspondence(db)The full pipeline from raw Statistics Canada inputs to verified artifacts uses package functions only. No source checkout is needed. Each step is a separate function call so you can inspect intermediate results.
The build requires sf (system libraries GDAL, GEOS,
PROJ), dplyr, and readr. On Ubuntu:
sudo apt-get install libgdal-dev libgeos-dev libproj-dev.
On macOS: brew install gdal geos proj. On Windows: install
Rtools and
use the sf binary from CRAN.
None of the build steps below are evaluated when this vignette is
built: they download the raw public inputs and need sf,
dplyr, and readr. Each step also consumes the
output of the step before it. Raw source files are large, so all
download helpers require an explicit directory that you manage.
packages <- c("digest", "dplyr", "jsonlite", "readr", "sf")
missing <- packages[!vapply(packages, requireNamespace, logical(1),
quietly = TRUE)]
if (length(missing)) install.packages(missing)Step 1: Download all public inputs. Each download is cached; re-running skips files already present.
build_cache <- file.path(tempdir(), "opcc-build")
dir.create(build_cache, recursive = TRUE, showWarnings = FALSE)
nar_dir <- download_nar(cache_dir = build_cache)
geonames_txt <- download_geonames(cache_dir = build_cache)
bounds <- download_census_boundaries(cache_dir = build_cache)
gaf_csv <- download_gaf(cache_dir = build_cache)Step 2: Build postal code centroids (M1). Reads NAR address points and GeoNames, computes the mean coordinate per postal code. NAR centroids take priority; GeoNames fills gaps.
output_dir is required: OPCC never picks a destination
for large build products on your behalf.
centroids_csv <- build_centroids(nar_dir, geonames_txt,
output_dir = file.path(build_cache, "centroids"))Step 3: Assign centroids to DBs and join GAF (M1). Validates Ontario boundary membership, assigns each centroid to one 2021 Dissemination Block via point-in-polygon, and joins the Geographic Attribute File for DAUID and higher geographies.
rollup_csv <- build_db_assignment(
centroids_csv, bounds$province, bounds$db, gaf_csv,
output_dir = file.path(build_cache, "rollup")
)Step 4: Build the DB correspondence (M2). Reads raw NAR address points, assigns them to DBs, aggregates evidence per postal-code/DB pair, and appends GeoNames supplementary links.
m2_csv <- build_m2(nar_dir, bounds$db, gaf_csv, rollup_csv,
output_dir = file.path(build_cache, "m2"))Step 5: Reproduce the DA roll-up (M5). This is a deterministic attribute roll-up of the M2 artifact. No spatial join or new input is needed.
Step 6: Verify.
stopifnot(!anyDuplicated(db[c("postal_code", "DBUID")]))
weights <- tapply(db$allocation_weight, db$postal_code, sum)
stopifnot(all(abs(weights - 1) < 1e-8))
stopifnot(all(tapply(db$best_link, db$postal_code, sum) == 1L))Exact byte-level rebuild. To reproduce the published
artifact bytes exactly, check out the producer revision recorded in the
published manifest rather than the tip of main. See docs/m2-reproduction.md
and docs/m5-reproduction.md
for the pinned revisions, expected SHA-256 values, and full verification
sequences.