Files
DeepSeek-TUI/telemetry-ingest
CodeWhale Bot b9dad36472 feat(telemetry): report observed active installs with trend, freshness, and caveats
Operationalize the owner report around the honest metric: observed active
installs = distinct rotating anonymous install ids with a session_start
ingested on a UTC day. The canonical report:active-installs command prints
the daily series, a complete-days 7-vs-7 trend, event freshness, and an
always-printed coverage-caveat block; report:dau remains as a pure re-export
compat alias. Exclusion guarantees are now pinned by tests (index1 only
inside count(DISTINCT), no content/identity/network columns anywhere in the
report path), and the output is forbidden from claiming DAU/unique users.

Known floor semantics recorded in docs: day attribution is ingest-day (events
carry no per-event timestamps), distinct counts cannot be sampling-corrected,
and id rotation can double-count across a trend window — all framed as the
lower-bound caveats they are.

Verified: telemetry-ingest npm test (109/109), npm run typecheck, CLI smoke
against fixtures.

Implemented with Claude Code agent assistance.
2026-08-10 14:15:04 -07:00
..

telemetry-ingest — Codewhale's first-party telemetry endpoint

A Cloudflare Worker that accepts the batches described in docs/TELEMETRY.md and writes them to Workers Analytics Engine. One POST route. No response body on any path. No client IP, anywhere, ever.

It lives here and not in web/ because the site is a separate deploy with its own build (Next.js via OpenNext); this is a single 13 KiB script with no assets, and coupling the two would mean a telemetry change rebuilding the marketing site.

Deployed and live at https://telemetry.codewhale.net/v1/telemetry, which is the shipped default for telemetry_endpoint. workers.dev is disabled; that hostname is the only way in.

Anonymous usage counting is on by default in v0.9.6, with a clear first-run disclosure and a durable opt-out. Prior declines remain off. A user who wants to contact nobody sets telemetry_endpoint = "", which writes batches to $CODEWHALE_HOME/telemetry/dryrun.jsonl instead.


The one property that matters

docs/TELEMETRY.md publishes:

Batches are IP-stripped at ingest. No IP is stored, logged, or joined to install_id.

This Worker is the whole of what makes that sentence true. There is no other component. So:

  • src/index.ts reads exactly two request headers — content-type and content-length — and nothing else, ever.
  • It never touches the cf property of the request, so country, colo, city, region, ASN and coordinates are never in scope.
  • src/datapoint.ts builds every stored row, and it cannot see the request at all: its input type is the validated batch body.
  • Nothing logs. invocation_logs is off in wrangler.jsonc, because Cloudflare describes those as "enriched with information available to Cloudflare in the context of the invocation" — exactly the class of automatic per-request record this service promises not to keep.
  • test/no-ip.test.ts reads the shipped source as text and fails the build if any of those names appears, if the set of headers read grows past two, if a console.* call is added, or if a Response is ever constructed with a body. That file is the only place in this directory where the forbidden header names are written down. A later edit cannot add one quietly.

Debugging without an IP is a solved problem: the schema carries os, arch, libc, surface, app_version and git_sha, which is what crash triage actually needs.

What it stores

Everything in docs/TELEMETRY.md and nothing else. The validator in src/schema.ts is a closed field set: an unexpected key anywhere in the batch — envelope, event, counters, errors, turn_wall — rejects the whole batch with 400. That is the point of the design. A future client bug that starts attaching a path, a prompt, or a customer's provider table name gets refused by the server rather than quietly stored.

test/schema-doc.test.ts parses the field names and enum spellings back out of docs/TELEMETRY.md and asserts set equality against the validator, and test/ingest.test.ts posts crates/telemetry/tests/golden/v1.json — the client's own pinned v1 wire form — and asserts it is accepted byte for byte and that deleting any key from it is rejected. The doc, the Rust client, and this endpoint cannot drift apart without a red test.

Column layout

One Analytics Engine data point per event. A batch carries at most 200 events (BATCH_MAX_EVENTS) and Analytics Engine allows 250 data points per invocation, so a conforming batch never needs a second pass.

The layout is positional and append-only — Analytics Engine columns are blob1..blob20 / double1..double20, and the names below exist only in the SQL you write. Renumbering silently rewrites every historical query. To add a field, take the next free slot.

column contents
index1 install_id — random v4 UUID, client-rotated every 90 days. The only identifier in the schema.
blob1 eventinstall_or_upgrade | session_start | session_end | panic
blob2 surface
blob3 os
blob4 arch
blob5 libc
blob6 app_version
blob7 git_sha, '' for every locally built binary
blob8 tty'true' | 'false'
blob9 install_kind (install_or_upgrade only)
blob10 previous_version (install_or_upgrade only)
blob11 session_source (session_start only)
blob12 duration_bucket (session_end only)
blob13 exit_class (session_end only)
blob14 cold_start_bucket (session_end only; '' on surfaces that do not measure it)
blob15 providers, comma-joined, already sorted and deduplicated
blob16 panic_site (panic only) — a crates/… path or the literal <dep>
blob17 sent_at — the batch timestamp. Events carry none.
double1..10 counters: turns, tool_calls, fleet_dispatch, workflow_run, subagent_spawn, mcp_server_connected, memory_search, approval_modal_shown, approval_auto_allowed, command_palette_open
double11..16 errors: auth_preflight_failed, provider_http_4xx, provider_http_5xx, tool_denied_by_policy, tool_timeout, network_error
double17..20 turn_wall: lt_5s, 5_30s, 30_120s, gte_120s

Columns not relevant to an event are '' / 0. tty is a blob because the 20 doubles are exactly used by the three numeric structs — Analytics Engine's ceiling is 20.

What it structurally cannot store

Not "does not"; cannot, given the code as written:

  • The client IP, and anything derived from it — never read. See above.
  • Any geo — country, colo, city, region, ASN, coordinates, timezone.
  • Any key the schema does not name. Unknown key ⇒ 400 for the entire batch, so there is no path from an unexpected field to storage.
  • Any free-form string. The published schema has no free-form string type and no open-keyed map. Every field is an integer, a boolean, or a closed enum, except app_version, git_sha and panic_site — each of which has a regex here, so a path, a prompt, a URL, or a branch name fails the shape check.
  • A provider table name. providers entries must be lowercase hyphenated ids and the array must be sorted and deduplicated; acme_internal_gateway is rejected. (See "known gap" below.)
  • A panic message. Only panic_site, and only inside the crates/ allowlist or the literal <dep>.
  • Per-event timestamps. There are none in the schema; only sent_at, per batch.
  • Response content. Every response is a bare status with a null body, so the endpoint cannot echo back what it received or what it holds.

Known gap, stated plainly. providers is the one field whose value space this endpoint cannot close. The authoritative list is codewhale_config::provider::all_providers(), a Rust registry with no generated artifact to read, and hard-coding a copy here would drift into silently dropping a real user's route. The client closes it (Event::is_boundedis_known_provider_id) before the POST is made; the server enforces the shape a closed &'static str enum can produce, plus the doc's sorted-and-deduplicated rule. If a generated provider-id list ever lands in the repo, wire it in here.

Retention

Cloudflare stores Analytics Engine data for three months, and that is not configurable — a ceiling rather than a policy, since no setting could make it longer. docs/TELEMETRY.md states it.


Deploy

Live. The commands below are the ones that produced the current deployment and the ones that will produce the next one.

cd telemetry-ingest
npm install
npm test                    # 109 tests, including the doc weld and the IP guard
npx wrangler deploy --dry-run --outdir=.wrangler/dry-run   # no account touched
npx wrangler deploy         # <- the only command that publishes anything

The routes block in wrangler.jsonc binds it to telemetry.codewhale.net as a custom domain, so the endpoint URL is https://telemetry.codewhale.net/v1/telemetry. The workers.dev subdomain is disabled: that hostname is the only way in.

Verified against the live endpoint before the client default was changed: the client's golden batch returns 204 with a zero-byte body; an unknown key returns 400; GET returns 405; a wrong content type returns 415; a POST to / returns 404. Reading back from Analytics Engine returned exactly two rows — session_start and session_end, carrying install_id, surface=tui, os=macos — the documented shape and nothing else.

Re-run the verification below after any deploy. The client default now points here, so a regression in this Worker is a regression in a promise docs/TELEMETRY.md makes to users.

Analytics Engine dataset setup

There is none. The dataset named in wrangler.jsonc (codewhale_telemetry) is created implicitly on the first successful writeDataPoint, so there is nothing to provision ahead of the deploy. Confirm it exists after the first batch:

curl "https://api.cloudflare.com/client/v4/accounts/$CF_ACCOUNT_ID/analytics_engine/sql" \
  --header "Authorization: Bearer $CF_API_TOKEN" \
  --data "SHOW TABLES"

The token needs Account → Account Analytics → Read. Querying is out of band through this API; the Worker itself has no read path at all.

Rate limiting

ratelimits binding, 20 POSTs per 60s, keyed on install_id from the validated batch body. Never on a network address — an IP-keyed limiter would mean this Worker handles IPs, which is the whole thing it must not do. That is a weaker limiter (an install_id.json can be rewritten between POSTs) and it is the right trade: Cloudflare's edge already absorbs volumetric abuse, and the client only flushes once per session anyway (one three-second attempt at shutdown).

Size cap

MAX_BODY_BYTES is 72 KiB, computed rather than guessed. A conforming client sends at most 200 events totalling at most 65536 bytes (BATCH_MAX_EVENTS and BATCH_MAX_BYTES in crates/telemetry/src/actor.rs, both hard ceilings because parse_events breaks before crossing them), plus 199 commas and ~375 bytes of envelope keys and values — 66110 bytes worst case. 72 KiB is ~11% headroom.

The 512-record / 256 KiB rings in crates/telemetry/src/buffer.rs are the disk cap, not the wire cap: a full ring drains as three batches, never as one POST.

content-length is checked first as a cheap reject, but it is client-supplied, so the real bound is enforced while reading the body and the stream is cancelled the moment it goes over.


Owner reports and queries

The routine activity report is a checked-in command rather than SQL copied from a chat:

CF_ACCOUNT_ID=... CF_API_TOKEN=... npm run report:active-installs
CF_ACCOUNT_ID=... CF_API_TOKEN=... npm run report:active-installs -- --days 30 --json

The metric is observed active installs by UTC day: distinct rotating anonymous install_id values that produced a session_start event that day, plus the weighted number of sessions started. It is never a count of people, accounts, or total installs, and the report says so next to the numbers. npm run report:dau still works — scripts/report-dau.mjs is a pure re-export of the canonical scripts/report-active-installs.mjs, kept only so the old name keeps working — but the canonical name matches what the number actually is.

One command produces all four things the ledger asks for:

  • the daily series (default --days 15 — 14 complete UTC days plus the partial current day, marked *);
  • a 7-day trend over complete UTC days only: the last 7 days' sum of daily observed active installs against the previous 7, with the percentage change. When the window cannot cover both sides it says so instead of printing zeros;
  • event freshness: the timestamp of the newest ingested event of any kind and how stale it is, or a plain statement that nothing has been ingested;
  • the coverage caveats, printed with the numbers in both text and --json modes: clients older than the telemetry feature, opted-out installs, and non-emitting environments (kill switches, fleet workers, offline shutdowns, dropped flushes) are invisible, so every count is a lower bound; and the 90-day id rotation means week-over-week comparisons are not a retention metric.

The token needs Account → Account Analytics → Read. The report performs two read-only SQL requests (the series and the freshness probe), does not print credentials, and rejects windows beyond Analytics Engine's fixed 90-day retention. Its read path is pinned by test/report-active-installs.test.ts: the install id appears only inside count(DISTINCT …), no payload column (blob2+, double*) is ever selected, and the output wording can never label the result as users, people, or accounts.

The underlying ad-hoc queries are one query each, which is what the column layout was chosen for. Run them against the SQL API:

query() {
  curl -sS "https://api.cloudflare.com/client/v4/accounts/$CF_ACCOUNT_ID/analytics_engine/sql" \
    --header "Authorization: Bearer $CF_API_TOKEN" --data "$1"
}

_sample_interval says how many original rows a stored row represents; Analytics Engine downsamples high-volume indexes, so every count is weighted by it rather than using bare count().

(a) How many installs and sessions

SELECT
  count(DISTINCT index1)                                          AS installs,
  sumIf(_sample_interval, blob1 = 'session_start')                AS sessions_started,
  sumIf(_sample_interval, blob1 = 'session_end')                  AS sessions_ended,
  sumIf(_sample_interval, blob9 = 'install')                      AS first_installs,
  sumIf(_sample_interval, blob9 = 'upgrade')                      AS upgrades,
  sumIf(_sample_interval, blob13 = 'clean')                       AS clean_exits
FROM codewhale_telemetry
WHERE timestamp > NOW() - INTERVAL '7' DAY

Read installs as what the published doc says it is and nothing more: the id rotates every 90 days and is regenerated whenever the telemetry directory is cleared, so no count derived from install_id is a user count. It is a lower bound on distinct machine-installs seen in the window, and it undercounts a returning user across a rotation.

Add , blob6 to SELECT and GROUP BY blob6 to cut by app_version; blob3 for OS, blob2 for surface.

(b) Which error classes and panic sites dominate

SELECT
  blob16                                   AS panic_site,
  sum(_sample_interval)                     AS rows,
  sum(double11 * _sample_interval)          AS auth_preflight_failed,
  sum(double12 * _sample_interval)          AS provider_http_4xx,
  sum(double13 * _sample_interval)          AS provider_http_5xx,
  sum(double14 * _sample_interval)          AS tool_denied_by_policy,
  sum(double15 * _sample_interval)          AS tool_timeout,
  sum(double16 * _sample_interval)          AS network_error
FROM codewhale_telemetry
WHERE timestamp > NOW() - INTERVAL '7' DAY
  AND blob1 IN ('session_end', 'panic')
GROUP BY panic_site
ORDER BY rows DESC

One query, two answers, because panic_site is '' on every non-panic row:

  • the panic_site = '' row carries the six error-class totals across all session_end events in the window — that is the error ranking;
  • every other row is one panic site, ranked by how often it fired.

Sessions that ended in a panic are visible either way: blob13 = 'panic' on the session_end row, and the panic event carries the site.


Verifying no IP is stored

Three checks, in increasing order of how convincing they are.

1. The source cannot ask for it. npm test runs test/no-ip.test.ts, which greps the shipped source. Prove the guard is live by adding a request.headers.get("CF-Connecting-IP") line to src/index.ts and re-running — three tests go red — then revert.

2. Every stored column is accounted for. The schema is closed and the layout above is exhaustive; there is no free slot an address could occupy. Confirm the deployed dataset has exactly the columns you expect:

query "SELECT * FROM codewhale_telemetry LIMIT 1 FORMAT JSON"

The result carries dataset, timestamp, _sample_interval, index1, and the blob*/double* columns. There is no address column, because Analytics Engine has no implicit one — every column is written by writeDataPoint.

3. Search the stored data for an address shape. After the first real batches, this returns zero rows:

SELECT count() AS suspicious
FROM codewhale_telemetry
WHERE timestamp > NOW() - INTERVAL '7' DAY
  AND (
    match(index1, '\\d+\\.\\d+\\.\\d+\\.\\d+')
    OR match(concat(blob1, blob2, blob3, blob4, blob5, blob6, blob7, blob8,
                    blob9, blob10, blob11, blob12, blob13, blob14, blob15,
                    blob16, blob17), '\\d{1,3}\\.\\d{1,3}\\.\\d{1,3}\\.\\d{1,3}|:[0-9a-f]{1,4}:')
  )

Also confirm nothing is being logged: with invocation_logs off and no console.* call in the source, Workers Logs → this Worker should show no per-request entries at all.


Local development

npm install
npm test          # vitest, 109 tests
npm run typecheck # tsc --noEmit
npm run check     # wrangler deploy --dry-run — touches no account
npm run dev       # wrangler dev --local

wrangler dev --local runs the real workerd with local Analytics Engine and rate-limit bindings. Post the client's own golden batch at it:

curl -i -X POST http://127.0.0.1:8787/v1/telemetry \
  -H 'content-type: application/json' \
  --data-binary @../crates/telemetry/tests/golden/v1.json
# HTTP/1.1 204 No Content, zero-byte body

curl -o /dev/null -w '%{http_code}\n' http://127.0.0.1:8787/v1/telemetry
# 405, Allow: POST

Responses

status when
204 accepted, zero-byte body, no headers
400 not JSON, or fails the published schema — including any unknown key
404 POST to a path other than /v1/telemetry
405 any method other than POST, on any path (Allow: POST)
413 body over MAX_BODY_BYTES
415 content type is not application/json
429 this install_id is over the rate limit
500 internal error; nothing was written

Every one of them has an empty body. The client (crates/telemetry/src/client.rs) reads only the status class and drops the batch on anything that is not 2xx — no retry, no backoff, no re-queue — so a rejection is invisible to the user by construction, and a 5xx here can never become a client-visible error. That is what lets this endpoint fail closed: when in doubt, refuse the batch.