The store-path index
This is every indexed (attribute, version) pair, matched to the store path Hydra
built for it, which is what makes fast.*
possible, and what the site's cache-liveness, dependency and closure views
draw from.
Two sources answer two different questions. Evaluating nixpkgs at a
revision and an explicit system says what the store path is; every
nixos-unstable channel bump's published listing (store-paths.xz, or a
MANIFEST for the 2012-2013 bumps) says whether Hydra built it. Join the
two and every historical version gets a concrete
/nix/store/<digest>-<name> that cache.nixos.org
can still serve.
A path belongs to one system
A listing is a flat list of paths with no System column, and it holds every
system that channel bump was built for. At the 2026-08-17 tip that is 219,623
paths over 106,909 distinct derivation names: 98.4% of names appear more than
once.
/nix/store/bf2z97cfq7x47y1dy50k04grkikmjcgg-hugo-0.164.0 <- aarch64-linux
/nix/store/zlyg48yqjd4jwz5yz3zhrnac81qgy7mh-hugo-0.164.0 <- x86_64-linux
So a listing cannot be asked "which digest goes with this name" โ there is no answer to that question, and the index used to take whichever path sorted first and hand x86*64 users aarch64 binaries about half the time (#12). It can only be asked *"is this exact digest one you hold"_, which is a question with an answer.
The artifacts are therefore per system: outpaths-x86_64-linux.json,
outpaths-aarch64-linux.json, outpaths-aarch64-darwin.json, and the same for
the tip and sibling-output files. A system with no artifacts is not served a
neighbour's โ fast.* throws and names the eval selector instead.
The listings come from nixos-unstable, so they hold Linux paths and nothing
else: at the 2026-08-17 tip they named 23 of aarch64-darwin's 16,545 evaluated
paths, against 93.6% of x86_64-linux's. Darwin is carried entirely by the cache
probe below, and its coverage starts in 2021 โ a 2019 nixpkgs cannot be
evaluated for aarch64-darwin at all.
The site follows the same rule. Its aggregate views โ reverse dependencies, the
census, the universe map โ are built from one system, and a package page shows
that system's store paths with a picker to switch between all of them, carried
in the URL as ?sys= so a link names the system it was read on. Each system's
metadata lives in its own meta-<system>/ shards, the aggregated one included,
and an alternate is fetched only when a reader picks it, so a page nobody
switches costs exactly what it did before.
Resolving a pair
For each (attribute, version) pair, at the newest revision that shipped it:
- evaluate that revision at this system and take the attribute's
outPath(nix/eval-outpaths.nixundernix-eval-jobs, one file per revision and system); - keep the digest if the revision's listing holds it;
- otherwise ask cache.nixos.org directly, since a listing describes one
evaluation while the cache is the union of every jobset Hydra ever ran โ
firefoxat the tip is in the cache and not in the listing, and every darwin path is in this class; - otherwise record a miss. Nothing is guessed: a pair with no proof carries no
entry, and
fast.*throws rather than substituting something plausible.
Across 2016 to the tip, 91.5-95.0% of the Linux systems' evaluated attributes
are in their revision's listing verbatim. The remainder is (a) attributes Hydra never built,
and (b) attributes this evaluation builds differently from Hydra, which sets
allowUnfree = false where the index sets it true โ hplipWithPlugin,
caffeWithCuda, _7zz-rar. Under 0.3% of attributes per revision.
Class (a) is wider than "unfree or broken". meta.hydraPlatforms = [ ]
takes an attribute out of the jobset, and wrapper packages use it routinely
so Hydra does not rebuild a symlink farm.
The digest is per version, not per revision
The index records one digest per version, taken at the newest revision that
shipped it. That is the build-correct choice: the most patched,
most recently built, most likely to still substitute, and it is why the
fast.* honesty classes read the way they do: exact-version selectors are
bit-exact, while revision selectors are version-exact but build-canonical
(the right version, as its newest build, which may come from a slightly
newer revision than the one named).
It is not per branch either, which is why releases does not work.
Everything here is joined against nixos-unstable listings,
and a (attribute, version) pair names a different build on every branch.
A branch may have a different stdenv, different patches, or different build flags, for a single package which bubbles up and changes the store path of nearly every package, see releases have no fast path.
The census
A matched digest is a claim that the path substitutes. The census re-earns
that claim: for every indexed digest, GET the narinfo and HEAD the NAR
payload it points at โ the cache remembering a path and the cache still
serving its bytes are different claims. The initial census verified 100.00%
of matched paths alive, down to every NAR payload file; a weekly workflow
(census.yml) repeats the sweep, publishes
the snapshot to the rolling release, and feeds the result back into the
artifacts so the site and fast.* stop advertising a path that is gone. It
writes both directions โ a path that answers again after being recorded dead
is written back as alive.
Absence is not death
Liveness has three states and info-indexed has one field to spell two of
them, so the third is spelled by leaving the digest out. An entry exists only
where something looked: a crawl record, a MANIFEST-era listing, or a verdict a
previous cut published. A digest with none of the three is not in the file โ
the stats page's matched drops it, a package page shows no badge, and
fast.* is unaffected either way, since it reads the outpaths files and never
this one.
Verdicts accumulate rather than expire. Each run carries forward every entry it has nothing newer to say about, so what the file claims for a digest is the last fetch anyone made against it, however long ago. That is what makes the weekly sweep's resurrections worth as much as its deaths, and why a row is only ever written from a fetch.
The crawl graph holds the same line: it records what has been fetched and
nothing else. A runner that restores no graph re-crawls โ about 2.3M narinfos,
a quarter of an hour, in a job that allows ninety minutes and normally
finishes in ninety seconds. checks.store-liveness covers the rule.
Multi-output packages
The listings record each derivation's default output. But consumers
reference the other outputs (ffmpeg-7.1-lib, ffmpeg-7.1-bin), so the
dependency crawl already fetched their narinfos; joining them back gives
every multi-output package its sibling outputs with sizes and references.
Fakes expose them (fast.latest.ffmpeg.lib), and the site lists them.
fast.* reads these from outs-<system>.json, where they are keyed by the
out path's digest. Keying them by derivation name โ as this file did before
the per-system split โ let a name claimed by two packages, or by two
architectures, hand back somebody else's lib.
The site publishes the same maps split by digest, at
outs-<system>/<xx>.json (tools/shard-outs.py): around a thousand shards of
a few KB per system, keyed as the artifact is:
{ "<out digest>": { "bin": "<digest>", "lib": "<digest>", "man": "<digest>" } }
Every suffix the join recorded is there, minus siblings that repeat the digest they are filed under. The meta shards carry each sibling's suffix and size; these carry its digest.
The evaluation reports every output of a derivation directly, so nothing here
depends on some consumer having referenced an output. The site's own
outs-indexed view still recovers siblings from closures, because it wants the
sizes and references the crawl collected; an output nothing referenced can be
missing there.
Where the data lives
Three tiers, decided by one question: does anything pin it?
- The repository tree keeps only what evaluation reads offline:
revisions.json,releases.json, the index files โ anddata-pins.json, which is the only thing evaluation-facing code ever sees of the store-path artifacts. - Dated releases (
data-YYYYMMDDtags on this repository) carry the pinned artifacts as assets:outpaths-<system>.jsonandouts-<system>.jsonwhole (the fast path does point lookups and fetches exactly one small file โ which is why the system is in the filename rather than a key inside each entry, so an x86_64 user never downloads the aarch64 half), the graph artifacts (info-indexed,refs-indexed,closures) sharded by each digest's closing period โ year files for finished years, month files for the current one โ so a cut re-uploads only the shards that moved. Assets on a dated tag are immutable by convention; the narHash in each pin fails closed if the convention is ever violated. Consumers fetch withbuiltins.fetchTree { type = "file"; ... }, lazily, keeping this flake'sinputs = { }founding line intact.data-pins.json.baseUrlis the canonical location;fetchArtifactpoints the fetch at a mirror, or at a local directory. The fetcher covers artifacts read by the multiverse API, not the publishing and restoration tools. - The rolling release (
data-rolling) carries the working state between cuts: the currentoutpaths-<system>.jsonandtip-outpaths-<system>.json, the census snapshots, the per-system miss lists, and the crawl graph the incremental jobs resume from. Every bump rewrites it; a dated cut freezes whatever it holds at the time.
A lagging pin is harmless by design: the delta between cuts is "versions that closed since" โ things that were current yesterday. A stale pin loses the zero-eval fast path for exactly those versions, and the eval fallback serves them meanwhile.
That property is why tip-outpaths-<system>.json is safe to pin at all, and why it
is read only as keyed data โ (attr, version) โ digest โ never as a
statement about which revision is current. The dated cut happens on the
first data run of each UTC day and is skipped for the rest of it, so the
snapshot's own revisionCount falls behind revisions.json within hours.
Selectors resolve against revisions.json; this file only answers "do you
have a digest for this exact pair". An artifact claiming more revisions
than revisions.json holds is refused outright, since it cannot be
describing the same history.
The pipeline
The hourly update-index workflow
appends a data pass after the index update, all of it incremental (the
scripts live in tools/, the evaluators they drive in nix/, and
update-outpaths.sh orchestrates them):
- fetch the new bump's listing (
fetch-store-paths.py); - evaluate the new revisions, once per published system
(
eval-outpaths.sh, twonix-eval-jobsworkers on a standard runner); - join the two into digests (
join-eval-listing.py: pairs that closed before the previous cut are carried over, only the delta is resolved); - crawl cache.nixos.org for the newly resolved digests and their transitive
references (
crawl-narinfos.py, resuming from the rolling crawl graph, and crawling the lot again if it restored none); - consolidate into the artifact files (
consolidate-outpaths.py), with the previously published copies as fallback for digests this runner never crawled; - recover the site's sibling-output view (
extract-outputs.py).
Once a day, the first run after a channel bump shards the artifacts by
closing period (shard-data.py), uploads whatever differs from its pin to a
dated tag, and repoints data-pins.json (cut-data-release.sh) โ the one
pin-churn commit a day. No bump, no cut.
The one-time backfill โ every listing, the full 1.4M-path crawl, and every revision evaluated for every published system โ runs on a big machine and seeds the dated release; CI never re-runs it. The evaluation half is the expensive one: 1,536 revisions ร 3 systems, at roughly 20 (revision, system) pairs a minute on a 256-core machine running 20 revisions at once. Adding a system is that run for the new system alone.
Adding a package set costs seconds, not hours
Each evaluation is cached as
index/.eval/<rev>.<system>.<evaluator>-<list>.json, and the two halves of
that key are there because the two inputs differ in what changing them can do.
nix/nested-sets.nix is read in exactly one branch of the evaluator, one that
fires only for the attribute it names, and that attribute reported nothing
before it was listed โ a package set is a non-derivation attrset, which
projects to {} and is neither reported nor recursed into. So adding a set can
only add rows. It cannot move or drop one, which means every row already on
disk is still exactly right.
nix/eval-outpaths.nix carries no such guarantee. Flip allowUnfree in the
config it hands nixpkgs and thousands of rows ought to vanish; a cached file
that kept them looks entirely plausible and names builds nobody makes.
So a list edit does not need the pipeline re-run. eval-outpaths.sh --topup
evaluates the listed sets alone and folds their rows into the evaluation
already on disk:
NIXPKGS=/path/to/nixpkgs nix run .#eval-outpaths -- --topup -j 40
It replaces the nested rows rather than appending to them, so removing a set
from the list drops the rows it used to contribute โ an append would leave
those behind forever. It refuses to run when the evaluator half of the key has
moved as well, since then the carried-over rows are the thing in question;
--assume-additive overrides that for a change you have checked yourself.
The difference is the whole reason the key has two halves. Measured on one
revision at x86_64-linux: 52 seconds to evaluate the top level, 3 seconds
to evaluate jetbrains alone, and the merged file is byte-identical to the one
the full pass produces. Across 1,541 revisions and 3 systems that is about 20
minutes against about 3 hours.
tools/merge-nested-eval.py does the fold and is where the correctness lives:
it refuses a base and a nested run that disagree about the revision or the
system, and refuses a nested run that reported a bare top-level name, since
either would quietly corrupt the file the join reads. checks.topup-merge
covers all three.
Credits
The fake-derivation technique that turns these digests into installables โ
build an attrset that walks like a derivation and let appendContext give
its outPath real store context โ is
tomberek's, from
fastpkgs. The store-path index is
what lets it cover every version back to 2013.