Waypoint Ledger

Integrity

How the count is kept honest

This tool asks the public to flag where a published federal figure is wrong. A count like that is only worth reading if the people reading it can check that nobody edited it afterwards, including us.

Check the chain yourself
  • SHA-256 chainEvery row is hashed together with the row before it.
  • One thumbper browser per figure.
  • 3 an hour, 8 a dayper network per figure, at most.
  • No accountA correction never requires one.

What it proves, and what it does not

It proves

that no row was edited, removed or reordered after it was written without the head changing. If we quietly deleted the corrections that were inconvenient, the head published here would no longer match the file you downloaded.

It does not prove

that every row came from a different person, or that we published every row we received. Those are different claims, and this page does not make them.

What the chain is

When a row arrives it is hashed together with the hash of the row before it:

row_hash = SHA-256( prev_hash + canonical_json(published fields) )

The first row of each chain uses 64 zeros. Canonical JSON means object keys sorted at every level, no whitespace, missing values written as null, so two people in two languages produce the same bytes. Both hashes are stored on the row and the current head is published below and at /api/integrity.

The heads, right now

Each chain below ends in one 64-character number. Change any published row, delete one, or move one, and that number changes. We walk every chain on request and print what we get, whether or not it matches.

Reading the heads and walking every chain…

What one vote costs

A demand signal anybody can move a hundred times is not a signal. A correction never requires an account, and nothing about an account is used to count one: the controls below have to work without knowing who anyone is, and they do.

An optional account exists for one purpose, opening your own saved ledger on another device, and a correction sent while signed in is counted exactly like an anonymous one.

Two controls, both stated plainly:

One thumb per browser per figure.

A browser sends a random identifier it keeps to itself; the server stores a one-way hash of that identifier, the figure and the table version, mixed with a secret only the server holds, and refuses a second thumb on the same figure.

Why the secret matters

Without it, anyone holding a browser identifier could recompute the stored value and learn whether that browser had spoken. The stored value is useless to everyone but this server, and it is never published or exported.

A cooldown per network per figure: 3 in an hour, 8 in a day.

Clearing a browser makes a fresh identifier, so the first control alone would be defeated by a person who cleared storage a hundred times. It is now three.

How the network is counted without being kept

The count is kept against a one-way hash of the network address with a daily salt and the same server secret, so it cannot be reversed into an address by anyone holding it, and it stops being comparable the next day. Nothing derived from it is ever stored in a published row.

What that does not stop

Someone determined, with a hundred browsers on a hundred networks. We will not pretend otherwise. What it costs them is now a hundred networks instead of a hundred clicks, and the register publishes the count on every figure and the days the corrections arrived, so the shape of any such push is visible in the numbers we publish rather than hidden inside them.

An honest person is never silenced by this

A refusal from the cooldown is a 429, and the browser keeps the correction and sends it when the window turns, so the fourth person in one household, or the fourth patient on one clinic’s wifi, loses nothing but a few minutes.

What each chain covers

Each row as it was written: the fields listed for each chain in the cards above and in the endpoint.

  • An optional note, the encrypted sentence on a survey, and every word written in an interview were never in the chain: we never publish them, so nobody outside could ever check a hash taken over them.

  • Two of the three open files now publish less than their chain covers, on purpose. The chain itself was not rewritten: every hash is the hash it was on the day the row arrived, taken over the full row.

survey.csv and gap.csv: what they leave out

survey.csv gives the day an answer arrived, never the time, and nothing the person said about themselves (age band, sex, coverage, region, state, stage, who answered, the ages at onset and diagnosis). One row carrying all of that could point at a person, so those answers are published only as counts, at /api/survey, and any answer fewer than 11 people gave is not shown there.

gap.csv gives the day, never the time.

Check it yourself

corrections.csv: recompute every hash

corrections.csv carries every field its chain covers, so you can recompute every hash from the file alone. Download it, walk it from the top, and recompute.

survey.csv and gap.csv: check the links

survey.csv and gap.csv leave out fields their hash was taken over, so from those files you check the links rather than recompute them.

The corrections recipe: the whole method, the same code we run (Python 3, no packages)

This is the whole method, the same code we run:

# Check our count without trusting us. Python 3, no packages.
import csv, json, hashlib, re, urllib.request

url = "https://waypointledger.org/api/export/corrections.csv"
req = urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0"})  # the site's edge refuses Python's default agent
lines = urllib.request.urlopen(req).read().decode().splitlines()
rows = list(csv.DictReader(l for l in lines if not l.startswith("#")))  # line 1 is a comment: the table version

def cell(s):    # a cell a spreadsheet would execute is stored with a leading
    if s[:1] == "'" and re.match(r"[=+\-@\t\r]", s[1:2]):   # apostrophe: strip it
        return s[1:]
    return s

def num(s):
    if s == "": return None
    return int(s) if "." not in s else float(s)

prev = "0" * 64
for r in rows:
    fields = {
        "received_at":   cell(r["received_at"]),
        "price_id":      cell(r["price_id"]),
        "verdict":       cell(r["verdict"]),
        "believed_usd":  num(r["believed_usd"]),
        "table_version": cell(r["table_version"]) or None,
    }
    canonical = json.dumps(fields, sort_keys=True, separators=(",", ":"))
    prev = hashlib.sha256((prev + canonical).encode()).hexdigest()
    assert prev == r["row_hash"], f'row {r["received_at"]} does not match'

print("head:", prev)  # compare with /api/integrity

One detail the code above handles: a cell that would otherwise begin with =, +, -, @, a tab or a carriage return is stored with a leading apostrophe, so a spreadsheet reads it as text instead of running it. Strip that one apostrophe before you hash. It is the only thing we do to a published value, and it is reversible.

The links recipe for survey.csv and gap.csv, and what it can and cannot show

The first prev_hash is 64 zeros, every later one is the row_hash of the row above, and the last row_hash and the number of rows match the head published above. That shows no row was dropped, added or moved since the head you compare with.

It cannot show that a published answer matches what was hashed; for that we walk every chain over the stored rows ourselves at /api/integrity?verify=1 and print the result, whether or not it matches.

# survey.csv and gap.csv: check the links row by row. Python 3, no packages.
import csv, urllib.request

for name, table in (("survey", "survey_responses"), ("gap", "gap_reports")):
    url = f"https://waypointledger.org/api/export/{name}.csv"
    req = urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0"})  # the edge refuses Python's default agent
    lines = urllib.request.urlopen(req).read().decode().splitlines()
    rows = list(csv.DictReader(l for l in lines if not l.startswith("#")))
    prev = "0" * 64
    for n, r in enumerate(rows, 1):
        assert r["prev_hash"] == prev, f"{name} row {n} does not link to the row above"
        prev = r["row_hash"]
    print(name, len(rows), "rows, head:", prev, "- compare with", table, "at /api/integrity")

The limit of a published hash

A row hash cannot be turned back into the row. It is still a fingerprint of the whole row, so someone who already knew every answer a particular person gave, and roughly when, could test that guess against it. That is the limit of what a published hash can hide, and we would rather say it than not.

The files, published oldest first, which is the order they were chained in:corrections.csvgap.csvsurvey.csv

What else is refused

  • A correction must name one of the published figures in the price table. An identifier that is not in the table is refused, with that sentence.

  • A second thumb on the same figure from the same browser is refused: you have already told us about this figure.

  • Free text is never exported, never served publicly, and never included in a share link.

  • Interview answers are encrypted before they are stored; the key is not in the code.

The method for the prices themselves is at how a number is made.Security contact and the vulnerability policy: security.txt.Accessibility: our conformance statement.