A backend service for ingesting, storing, and querying test evidence from heterogeneous sources (Bazel test logs, CI pipelines, manual test runs, HiL/PiL/vehicle tests).
Evidence Store provides a unified API to collect and query test results across different tools and workflows. It supports batch ingestion, cursor-based pagination, evidence inheritance across commits, configurable retention policies, and a web UI for manual test entry and search with regex filtering.
# Start Postgres + the server on :8000
docker compose up -d
curl http://localhost:8000/healthz
# Open the web UI
open http://localhost:8000
# In this repo: run our own tests and upload the results through the adapter
./scripts/dogfood.sh
To upload results from any other Bazel workspace, add the adapter as a
dependency and run it after your tests, the same way CI or dogfood.sh does:
# MODULE.bazel — published to the Bazel Central Registry, so this is all it takes
bazel_dep(name = "evidence_store_bazel", version = "0.0.3")
bazel test //...
bazel run @evidence_store_bazel//cmd/evidence-bazel -- \
--api-url http://localhost:8000 \
--testlogs-dir "$(bazel info bazel-testlogs)"
See Bazel adapter below for CI integration (GitHub Actions) and an always-on watch mode for a developer's own machine.
The rest of this document explains those commands and everything around them.
docker compose up -d
curl http://localhost:8000/healthz
This starts PostgreSQL 16, the Evidence Store server on port 8000, and (for the
s3 blob backend) a local MinIO. The web UI is served from the same port —
open http://localhost:8000.
| Variable | Default | Description |
|---|---|---|
EVIDENCE_DATABASE_URL |
postgres://evidence:evidence@localhost:5432/evidence_store?sslmode=disable |
PostgreSQL connection string |
EVIDENCE_LISTEN_ADDR |
:8000 |
Listen address |
EVIDENCE_LOG_LEVEL |
INFO |
Log level |
EVIDENCE_DEFAULT_PAGE_SIZE |
100 |
Default page size |
EVIDENCE_MAX_PAGE_SIZE |
1000 |
Max page size |
EVIDENCE_MAX_BATCH_SIZE |
1000 |
Max records per batch |
EVIDENCE_ANALYTICS_CACHE_TTL_SECONDS |
30 |
How long an analytics aggregation is reused for an identical filter (0 disables) |
EVIDENCE_QUERY_TIMEOUT_SECONDS |
15 |
Budget for one evidence query — a search or an aggregation — before it is refused (0 disables). EVIDENCE_ANALYTICS_QUERY_TIMEOUT_SECONDS is still read and supplies the default, from when the budget covered analytics only |
EVIDENCE_API_KEYS |
(empty — auth disabled) | Comma-separated API keys (see Authentication) |
EVIDENCE_AUTH_DB |
false |
Authenticate against the principals table (see Database-backed principals) |
EVIDENCE_BOOTSTRAP_ADMIN |
(empty) | Subject of an administrator seeded on first start; its key is logged once |
EVIDENCE_OIDC_ISSUER |
(empty — SSO off) | Identity provider to log people in with (see Single sign-on) |
EVIDENCE_OIDC_CLIENT_ID |
(empty) | Client id registered with the provider |
EVIDENCE_OIDC_CLIENT_SECRET |
(empty) | Client secret, for a confidential client |
EVIDENCE_OIDC_REDIRECT_URL |
(empty) | Where the provider sends the browser back, e.g. https://evidence.example.com/auth/callback |
EVIDENCE_OIDC_POST_LOGOUT_URL |
(the store's root) | Where the provider returns the browser after logging out; derived from the redirect URL unless set |
EVIDENCE_OIDC_PROVIDER_LOGOUT |
false |
End the provider's session on logout too. Off because "log out" usually means this application; worth turning on where a machine is shared |
EVIDENCE_OIDC_SCOPES |
openid,profile,email |
Scopes to request |
EVIDENCE_OIDC_GROUPS_CLAIM |
groups |
Claim carrying group membership (Entra calls it roles) |
EVIDENCE_GROUP_ROLE_MAP |
(empty) | group:role pairs for either provider, e.g. eng-all:contributor,eng-leads:admin. Also read by SCIM provisioning, so a group means one thing however the store hears about it |
EVIDENCE_SAML_IDP_METADATA_URL |
(empty — SAML off) | Identity provider metadata to fetch at startup (see SAML) |
EVIDENCE_SAML_IDP_METADATA_FILE |
(empty) | The same metadata from a file, for a deployment that will not reach out |
EVIDENCE_SAML_ROOT_URL |
(empty) | This store's public address, e.g. https://evidence.example.com |
EVIDENCE_SAML_ENTITY_ID |
(metadata URL) | What this service provider calls itself |
EVIDENCE_SAML_CERT_FILE |
(empty) | Service provider certificate, PEM |
EVIDENCE_SAML_KEY_FILE |
(empty) | Service provider private key, PEM |
EVIDENCE_SAML_EMAIL_ATTRIBUTE |
email |
Assertion attribute carrying the address |
EVIDENCE_SAML_NAME_ATTRIBUTE |
displayName |
Assertion attribute carrying the display name |
EVIDENCE_SAML_GROUPS_ATTRIBUTE |
groups |
Assertion attribute carrying group membership |
EVIDENCE_SESSION_TTL_HOURS |
12 |
How long a login lasts |
EVIDENCE_COOKIE_SECURE |
true |
false only for local development over plain HTTP |
EVIDENCE_RATE_LIMIT_READ_RPS |
0 (disabled) |
Sustained reads per second per caller (see Rate limiting) |
EVIDENCE_RATE_LIMIT_WRITE_RPS |
0 (disabled) |
Sustained writes per second per caller |
EVIDENCE_RATE_LIMIT_READ_BURST |
2 × read RPS |
Token-bucket burst capacity for reads |
EVIDENCE_RATE_LIMIT_WRITE_BURST |
2 × write RPS |
Token-bucket burst capacity for writes |
EVIDENCE_BLOB_BACKEND |
fs |
Where images live: fs or s3 (see Images in test logs) |
EVIDENCE_BLOB_PATH |
blobs |
Directory for the fs backend |
EVIDENCE_BLOB_S3_ENDPOINT |
(empty) | host:port of the S3/MinIO endpoint |
EVIDENCE_BLOB_S3_BUCKET |
evidence-blobs |
Bucket to store blobs in |
EVIDENCE_BLOB_S3_ACCESS_KEY |
(empty) | S3 access key |
EVIDENCE_BLOB_S3_SECRET_KEY |
(empty) | S3 secret key |
EVIDENCE_BLOB_S3_USE_SSL |
false |
true to talk to the endpoint over HTTPS |
EVIDENCE_BLOB_S3_REGION |
(empty) | S3 region |
EVIDENCE_MAX_BLOB_BYTES |
5242880 (5 MiB) |
Largest image that may be uploaded |
EVIDENCE_BLOB_ORPHAN_GRACE_HOURS |
24 |
How long an unreferenced image is kept before the sweep removes it |
EVIDENCE_RETENTION_CONFIG |
(empty — retention off) | Path to a retention rules YAML file (see Retention) |
EVIDENCE_WEATHER_ENDPOINT |
https://api.open-meteo.com/v1/forecast |
Forecast API the weather lookup asks. Set it to an empty value to switch the lookup off (see Weather while a test ran) |
EVIDENCE_WEATHER_TIMEOUT_SECONDS |
10 |
Budget for one weather lookup before the tester is told to type the conditions in |
Note that limiters live in memory, held under a digest of the caller rather than under the API key they authenticated with. Once a bucket holds more than a thousand callers, the ones whose token bucket has refilled are dropped — which cannot forgive anybody's rate debt, since a limiter created fresh starts full and a caller who still owes has a bucket that is not. Memory is therefore bounded by how many callers are active, not by how many have ever appeared.
Set EVIDENCE_API_KEYS to enable API key authentication for all /api/v1/* endpoints. The /healthz endpoint and static web UI files are always public.
Each key entry has the format role:key where role is rw (read-write) or ro (read-only):
# Single read-write key
export EVIDENCE_API_KEYS="rw:my-secret-key"
# Multiple keys with different roles
export EVIDENCE_API_KEYS="rw:ingest-key-for-ci,ro:dashboard-viewer-key"
rw keys can read and write (GET + POST).ro keys can only read (GET). POST requests return 403 Forbidden.401 Unauthorized.EVIDENCE_API_KEYS is empty or unset, authentication is disabled (open access).Behind those two key roles the server authorizes by permission, not by HTTP method. Authentication resolves the caller to a principal holding one or more roles, and every route states the permission it needs:
| Role | Permissions |
|---|---|
viewer |
evidence:read, analytics:read, blob:read, inheritance:read |
contributor |
viewer + evidence:write, blob:write |
ci |
contributor + source:any (may write a source that is not its own name) |
admin |
contributor + inheritance:write, principal:admin, retention:admin, scim:provision |
provisioner |
scim:provision only — a directory's own token, which reads nothing at all |
POST /inheritance requires inheritance:write, which only admin holds —
declaring that one commit inherits another's evidence is the elevated operation
DESIGN.md section 8 has always specified, and it used to be
indistinguishable from posting a test result.
Configured keys map onto those roles so that nothing that worked before stops
working: ro becomes viewer, and rw becomes ci and admin, since an
rw key can reach every endpoint today. To grant the finer roles individually,
give keys names and owners, or revoke one without a redeploy, use
database-backed principals below.
Clients authenticate by sending the key as a Bearer token:
curl -H "Authorization: Bearer my-secret-key" \
http://localhost:8000/api/v1/evidence
The Bazel adapter supports this via --api-key or EVIDENCE_STORE_API_KEY (see Bazel adapter). The web UI prompts for a key on first 401 and stores it in localStorage.
An entry in EVIDENCE_API_KEYS is a shared secret: no name, no owner, no
expiry, and no way to revoke one short of a redeploy. Setting
EVIDENCE_AUTH_DB=true turns on the principals table instead, where a key
belongs to somebody, holds exactly the roles it was granted, and stops working
on the next request when it is disabled.
| Variable | Default | Description |
|---|---|---|
EVIDENCE_AUTH_DB |
false |
true to authenticate bearer tokens against the principals table |
EVIDENCE_BOOTSTRAP_ADMIN |
(empty) | Subject of an administrator seeded on first start, e.g. user:ops@example.com. Requires EVIDENCE_AUTH_DB=true |
Switching it on closes the API. An empty principals table means nobody may
in, not that everybody may — so seed the first administrator at the same time:
export EVIDENCE_AUTH_DB=true
export EVIDENCE_BOOTSTRAP_ADMIN="user:ops@example.com"
On first start the server mints that principal a key and logs it once:
WARN bootstrap admin API key issued - copy it now, it is not stored and will not
be shown again subject=user:ops@example.com api_key=evs_...
Only a SHA-256 digest of the key is stored, so that line is the single moment it can be read. Miss it and the remedy is to rotate the key (see below), which keeps the identity and its roles. Restarts are safe: an existing subject is left alone and no second key is minted.
Keys are always minted by the server, never chosen by a caller — 256 bits from
crypto/rand, prefixed evs_. That is what makes a fast digest the right one
to store; see the comment on principals.key_hash in
migration 000006.
Both key sources run at once, which is the migration path: leave
EVIDENCE_API_KEYS in place, issue database keys, move pipelines over one at a
time, and clear the variable when the last has moved. Environment keys are
checked first because they cost no round trip. A token neither source
recognises is 401; if the database cannot be reached at all, requests get
503 rather than being told their key is wrong.
The Admin tab in the web UI is the everyday way in: it lists every
principal with its roles, when its key was last used, and whether it has been
revoked, and it issues, rotates and revokes keys. The tab is only shown to a
caller holding principal:admin. See Managing access for
what the tab looks like.
The same operations are /api/v1/principals, all behind principal:admin:
| Request | Does |
|---|---|
GET /api/v1/principals |
List every principal, revoked ones included |
POST /api/v1/principals |
Create one and mint its key — {"subject": "ci:nightly", "display_name": "…", "roles": ["ci"]} |
PUT /api/v1/principals/{id}/roles |
Set roles to exactly {"roles": [...]} |
POST /api/v1/principals/{id}/disable |
Revoke: the key stops working on the next request |
POST /api/v1/principals/{id}/enable |
Restore, with the roles it already had |
POST /api/v1/principals/{id}/rotate |
Issue a fresh key and invalidate the old one |
Creating and rotating return {"principal": {...}, "api_key": "evs_..."}. That
response is the only time the key can be read — only its digest is stored — so
a mislaid key is fixed by rotating, not by looking it up.
There is no delete. Revoking is a timestamp so that evidence already attributed to a principal still names something a reader can look up, and so an administrator can tell a credential that was taken away from one that never existed.
The last enabled administrator cannot be revoked or demoted; the request is
refused with 409 and a message saying to grant admin to somebody else first.
Otherwise one click could leave a deployment with no way in but psql.
GET /api/v1/me reports the calling principal, its roles and its permissions.
It is the one route under /api/v1 that asserts no permission of its own — the
web UI uses it to decide what to offer.
Point the store at your identity provider and people log in with the account
they already have. Set EVIDENCE_OIDC_ISSUER and the client credentials it
issued you:
export EVIDENCE_AUTH_DB=true # sessions resolve to principals
export EVIDENCE_OIDC_ISSUER="https://login.example.com/realms/engineering"
export EVIDENCE_OIDC_CLIENT_ID="evidence-store"
export EVIDENCE_OIDC_CLIENT_SECRET="…"
export EVIDENCE_OIDC_REDIRECT_URL="https://evidence.example.com/auth/callback"
export EVIDENCE_GROUP_ROLE_MAP="eng-all:contributor,eng-leads:admin"
Register the redirect URL with the provider, and make sure the ID token carries a groups claim — in Keycloak that is a group membership mapper, and most providers need it switched on explicitly.
To try this without a company directory behind it, there is a Keycloak in
docker-compose.sso.yml seeded with users whose groups map to each role — and
one whose group maps to nothing. See dev/keycloak/README.md.
The flow is Authorization Code with PKCE. GET /auth/login sends the browser
out, GET /auth/callback verifies the ID token and starts a session, and
POST /auth/logout ends it. The session is a row, not a signed cookie, so
revoking somebody stops the browser they left open rather than waiting for a
token to expire. GET /auth/config reports whether a login flow exists, which
is how the UI knows to offer one.
Logging out ends this session, and by default no other. /auth/logout
deletes the session row and answers with a logout_url for the browser to
follow — the store's own signed-out page, marked so the page knows this 401 is
what logging out looks like rather than an expired session to bounce back to the
provider.
Switch user sits beside Log in on the signed-out page, for the other half of
that: because the provider still knows who you are, an ordinary Log in is
answered instantly as whoever was here last. It sends /auth/login?switch_user=1,
which asks the provider to establish who this is rather than answer from the
session it holds.
That is spelled the store's own way rather than passed through as an OIDC
prompt, because which value achieves it differs by provider — Keycloak ignores
select_account outright and waves the browser through, while prompt=login is
honoured by both it and Entra. Demanding re-authentication is what switching
user means in any case: you have to prove you are the other person, not merely
name them.
Set EVIDENCE_OIDC_PROVIDER_LOGOUT=true to end the provider's session as well,
by sending the browser to its end_session_endpoint with the id_token_hint
from that login. It is off by default because it is a large side effect: the
person is signed out of every other application that account opens, and against
a real Entra tenant they are additionally asked to pick which identity they
meant to abandon. Turn it on where a machine is shared, since otherwise the next
person's Log in is answered silently as the last one.
With it on, register EVIDENCE_OIDC_POST_LOGOUT_URL with the provider alongside
the redirect URL; most refuse a post-logout redirect they were not told about. A
provider advertising no logout endpoint is fine — the session here still ends.
Roles come from groups. A group with no entry in EVIDENCE_GROUP_ROLE_MAP
grants nothing, so pointing this store at a company directory does not hand
every employee an account that can write. On each login the roles derived from
group claims are reconciled to what the token now says — losing a group loses
the role — while roles an administrator granted in the Admin tab are left
alone. Someone whose groups map to nothing is authenticated and permitted
nothing, which is a deliberate state and not an error.
People are matched on the provider's sub claim, not their address. Someone
who changes their email stays one principal with their history intact; their
subject here is corrected at the next login. If an API key already answers to
the name a login wants, the login is refused with 409 rather than guessing
they are the same party — rename the key.
Writes need a CSRF token. A cookie is sent by the browser whether or not the
page meant to send it, so session-authenticated writes must echo the
evidence_csrf cookie in an X-CSRF-Token header. Bearer-token callers are
unaffected: CI has no cookies and needs none of this.
Signing in also settles the source question for humans. A logged-in person's
Source box is filled in with their own subject and locked, because that is the
only value the server will accept from them — see
source is bound to the caller.
Same idea, different protocol, and the same everything else — a SAML login
produces the identical principal, session, roles and source binding an OIDC
one does.
export EVIDENCE_AUTH_DB=true
export EVIDENCE_SAML_IDP_METADATA_URL="https://login.example.com/app/xxx/sso/saml/metadata"
export EVIDENCE_SAML_ROOT_URL="https://evidence.example.com"
export EVIDENCE_SAML_CERT_FILE=/etc/evidence/saml.crt
export EVIDENCE_SAML_KEY_FILE=/etc/evidence/saml.key
export EVIDENCE_GROUP_ROLE_MAP="eng-all:contributor,eng-leads:admin"
A service provider needs its own X.509 keypair, which most identity providers will not register one without. A self-signed pair is enough:
openssl req -x509 -newkey rsa:2048 -nodes -days 3650 \
-keyout saml.key -out saml.crt -subj "/CN=evidence.example.com"
Then point the provider at https://evidence.example.com/auth/saml/metadata,
which describes this store in the XML they expect — where assertions go and
which certificate signs requests. Serving it beats writing it by hand, which is
how the two ends end up disagreeing about a URL.
GET /auth/saml/login starts a login and POST /auth/saml/acs is the Assertion
Consumer Service the provider posts back to. Logging out is the same
POST /auth/logout either way.
Attribute names vary a great deal between providers, so the three that matter
are configurable, and the common spellings — mail, cn, memberOf, and the
urn:oid: and schemas.xmlsoap.org forms Entra and ADFS send — are tried
automatically when the configured name is not present.
Both front ends can run at once. A company moving between protocols will have a period where each is somebody's way in, so the UI offers a choice when two are configured and goes straight there when there is one.
One detail worth knowing if you are reading the schema: an assertion arrives as
a form POST from the provider's own origin, and a SameSite=Lax cookie is
exactly what a browser will not send on a cross-site POST. The id of the request
it answers therefore lives in a saml_requests row rather than a cookie —
loosening the session cookie to SameSite=None to avoid that table would have
traded a real protection for a detail of one flow.
See docs/rbac-design.md for the whole plan, including where SSO/SAML plugs in.
source is bound to the callerA record's source is what a reader goes on months later to ask who ran this
and whether to believe them, so the server decides what it may say:
| Caller | source |
|---|---|
ci (holds source:any) |
Taken as sent — a build robot's useful attribution is the build URL, not the robot |
| Any other principal | Must equal the caller's subject. Left empty, the server fills it in; anything else is 403 |
| No principal (nothing configured) | Unchanged — there is no identity to pin it to |
In a batch, one record with a source the caller may not write refuses the whole
batch. A malformed record still behaves as it always did: reported in the
per-record results, with the rest of the batch filed and a 207.
admin does not subsume ci, so an administrator is pinned to their own name
too. Backfilling evidence in somebody else's name means holding both roles, so
that writing history under another party's identity is always a deliberate
grant. Configured EVIDENCE_API_KEYS are unaffected: every rw key is ci, so
the pipelines using one keep writing exactly the source they wrote before.
Single sign-on gets people in. It does not get them out.
A principal is created the first time somebody logs in, and their roles are reconciled from the group claims in that login's token. Both halves wait on a login happening — so somebody who leaves the company keeps their account here, and any browser session they left open, because the login that would have been refused never comes. A group change is revocation only if they sign in again. And nobody exists before their first login, so a joiner cannot be granted anything, or even seen, before their first day.
SCIM 2.0 is how a directory says all of this without waiting for the person: it calls this store on a schedule to create, update, deactivate and delete users and groups. It complements the OIDC login and replaces none of it.
export EVIDENCE_AUTH_DB=true # provisioned people are rows in `principals`
export EVIDENCE_GROUP_ROLE_MAP="eng-all:contributor,eng-leads:admin"
There is no switch to turn it on. The endpoints are always mounted and always
require scim:provision, so a deployment that has issued no token holding it
has no provisioning.
Mint the directory a token. In the Admin tab, create a key with the
provisioner role and nothing else, and give it to the directory as its secret
token. That role grants scim:provision and no reading of any kind — this is
the one credential in the store that lives for years inside another company's
configuration, and if it leaks the damage should be account creation, not a
customer's test results. admin also holds the permission, so an administrator
can drive the endpoints by hand.
| Endpoint | What it is for |
|---|---|
GET /scim/v2/ServiceProviderConfig |
Which optional parts of SCIM this store answers |
GET /scim/v2/ResourceTypes, /Schemas |
What can be provisioned, and which attributes are kept |
GET/POST /scim/v2/Users, GET/PUT/PATCH/DELETE /scim/v2/Users/{id} |
People |
GET/POST /scim/v2/Groups, GET/PUT/PATCH/DELETE /scim/v2/Groups/{id} |
Groups, and so roles |
Filtering is userName eq, externalId eq and displayName eq — the forms a
directory actually sends. Anything else is refused rather than answered with the
whole store, which to a client would look like every user matching every query.
Paging is startIndex and count, capped at 500.
Deactivating somebody ends their sessions. PATCH with active: false, and
DELETE, both disable the account and delete every session it has open, in
one transaction. Disabling alone would leave the browser they walked away from
working until it expired on its own, which is the failure provisioning exists to
close.
DELETE never removes the row. Evidence names its source, so a deleted
principal would leave records attributed to nothing, and an administrator could
no longer tell somebody who was revoked from somebody who never existed.
Roles come from groups, through the same EVIDENCE_GROUP_ROLE_MAP the login
path uses — one question, answered once. A group with no entry grants nothing,
so pointing this store at a company directory does not hand every employee an
account that can write. Removing somebody from a group takes the role away,
which is the only way a role comes back from somebody who keeps their account.
Grants are recorded with source = 'scim', alongside local (an administrator
granted it by hand) and idp (a login derived it from its own token). What
somebody may do is the union of the three, and each is reconciled only
against its own source — otherwise a login and a sync would delete each other's
grants on alternate runs.
The store will refuse to deactivate the last enabled administrator, answering
403 rather than leaving the deployment with no way in but psql.
A provisioned user and that same human's later login have to land on one principal. If they do not, everybody becomes two rows: one holding their evidence, and one that is the only thing the directory can deactivate — so deprovisioning would appear to work while the account that matters stayed live.
They cannot be matched on the obvious key. A login is identified by
<issuer>|<sub>, and Entra's sub is pairwise: unique to one application and
invented the first time somebody signs into it, so SCIM has never seen it and
cannot send it. The externalId SCIM does carry is mapped from mailNickname
by default and is configurable per tenant.
So a provisioned row starts unclaimed, and the first login that recognises itself in one takes it, writing its own external id there; from then on it is an ordinary login matched on that id. It works the other way round too: provisioning somebody who has already logged in adopts the principal they already have, which is what lets a deployment that had single sign-on first start managing the people who got there before the directory did.
This means your externalId mapping does not have to match anything here.
Both directions key off the address and login name instead, which is also why
the emails array matters: the primary address becomes the subject their
evidence is filed under, and the one a later login is recognised by.
Register the store as an enterprise application, then under Provisioning:
| Setting | Value |
|---|---|
| Tenant URL | https://evidence.example.com/scim/v2 |
| Secret Token | the provisioner key from the Admin tab |
Test Connection exercises the discovery endpoints and a probe query, so it fails immediately and legibly if the token or URL is wrong.
The default attribute mappings need one change: under Provision Azure AD
Users, make sure mail or userPrincipalName flows to emails[type eq "work"].value, since that address is what a later login is matched on. The
externalId mapping can be left at whatever your tenant uses.
Group provisioning is worth enabling: it is what makes membership decide roles
here rather than requiring an administrator to grant them by hand. Map the
groups you named in EVIDENCE_GROUP_ROLE_MAP.
Entra syncs roughly every 40 minutes, so a departure takes effect on the next cycle rather than instantly. Their open session dies the moment it does.
Set EVIDENCE_RATE_LIMIT_READ_RPS and/or EVIDENCE_RATE_LIMIT_WRITE_RPS to enforce per-caller token-bucket limits on /api/v1/* endpoints. Reads (GET/HEAD/OPTIONS) and writes (POST/PUT/PATCH/DELETE) use separate buckets so a flood of reads cannot starve writes.
# 50 reads/sec sustained (burst 100), 10 writes/sec sustained (burst 20).
export EVIDENCE_RATE_LIMIT_READ_RPS=50
export EVIDENCE_RATE_LIMIT_WRITE_RPS=10
X-Forwarded-For honored).429 Too Many Requests with a Retry-After header (seconds).Login sessions and half-finished SAML handshakes both carry an expiry and are checked against it on every use, so an expired row is already inert — but it still occupies a table. The server deletes them hourly, and does so unconditionally: unlike evidence retention below, this is not a policy an operator might reasonably answer "never" to, and the sweep costs nothing against the empty tables of a deployment that has no SSO configured.
Nothing to configure. It logs only when it actually deleted something.
Old evidence can be evicted automatically instead of growing the database
forever. Point EVIDENCE_RETENTION_CONFIG at a YAML file and a background
worker applies it on an interval:
export EVIDENCE_RETENTION_CONFIG=/etc/evidence/retention.yaml
See retention.example.yaml for a starting point —
rules match evidence by regex on fields like branch or evidence_type and
give each a max_age (0s keeps it forever), evaluated highest-priority
first. Records with metadata.retain=true, or referenced by an active
inheritance declaration, are always
exempt. Deleting a record releases the images its log referenced (see
Images in test logs). The full
rule syntax and rationale are in
docs/retention-rules-plan.md. Retention runs
default off — an unset variable means nothing is ever evicted.
Scans bazel-testlogs/ after a test run and uploads results to the Evidence
Store. It lives in adapters/bazel/ as its own Bzlmod module named
evidence_store_bazel, so other Bazel workspaces can depend on it without
pulling in the server's own dependencies.
The module is published to the Bazel Central Registry, so a plain bazel_dep in the consuming repo's MODULE.bazel is all it takes:
bazel_dep(name = "evidence_store_bazel", version = "0.0.3")
Adapter releases before 0.0.3 submit records with evidence_type: "bazel",
which the server has rejected since evidence types were restricted to ci,
manual_test, and demonstration (the runner now goes in
metadata.collector). Consumers of 0.0.2 or older see every upload fail
validation; upgrade to 0.0.3 or newer.
To track an unreleased commit instead of a published version — testing a fix
before it's tagged, say — add a git_override:
git_override(
module_name = "evidence_store_bazel",
remote = "https://github.com/nesono/evidence_store.git",
commit = "<pinned-sha>",
strip_prefix = "adapters/bazel",
)
For local development against a checkout instead:
local_path_override(
module_name = "evidence_store_bazel",
path = "/path/to/evidence_store/adapters/bazel",
)
Building it inside this repo (rather than as a dependency) is:
cd adapters/bazel
bazel build //cmd/evidence-bazel
A build-engineer setup: run the suite, then upload the results — pass or fail — so a broken run still leaves a record. A GitHub Actions job for a consuming workspace looks like this:
name: CI
on: [push, pull_request]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
- uses: bazel-contrib/setup-bazel@0.19.0
- name: Run tests
run: bazel test //... || true # keep going so results still upload
- name: Upload results to Evidence Store
env:
EVIDENCE_STORE_API_KEY: ${{ secrets.EVIDENCE_STORE_API_KEY }}
run: |
bazel run @evidence_store_bazel//cmd/evidence-bazel -- \
--api-url https://evidence.example.com \
--testlogs-dir "$(bazel info bazel-testlogs)" \
--source "${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}" \
--tags ci,github-actions
--repo, --branch and --rcs-ref are auto-detected from the git checkout, so
they are left out above. Mint EVIDENCE_STORE_API_KEY as a ci-role principal
(see Authentication and authorization) — it
needs source:any to attribute records to the build URL rather than to itself.
For a one-off run, bazel test then feed the adapter the resulting logs:
bazel test //...
bazel run //cmd/evidence-bazel -- \
--api-url http://localhost:8000 \
--testlogs-dir "$(bazel info bazel-testlogs)"
Or, from inside this repo, the dogfood script that does both:
./scripts/dogfood.sh
| Flag | Default | Description |
|---|---|---|
--api-url |
$EVIDENCE_STORE_URL |
API base URL (required) |
--testlogs-dir |
bazel-testlogs |
Path to testlogs directory |
--repo |
auto-detected | Repository (from git remote) |
--branch |
auto-detected | Branch (from git) |
--rcs-ref |
auto-detected | Commit hash (from git HEAD) |
--source |
auto-detected | CI build URL or username |
--invocation-id |
Bazel invocation ID | |
--tags |
Comma-separated tags | |
--api-key |
$EVIDENCE_STORE_API_KEY |
API key |
--dry-run |
false |
Print records instead of posting |
For continuous, hands-off uploads on your workstation — no changes to your
bazel test habit needed — run the adapter as a background watcher instead:
# One-time setup: create .evidence/config.yaml in your workspace
mkdir -p .evidence
cat > .evidence/config.yaml <<EOF
api_url: https://evidence.mycompany.com
tags: [local, dev]
EOF
# Start the watcher (runs in background)
bazel run @evidence_store_bazel//cmd/evidence-bazel -- watch start
# Check status
bazel run @evidence_store_bazel//cmd/evidence-bazel -- watch status
# Stop
bazel run @evidence_store_bazel//cmd/evidence-bazel -- watch stop
The watcher polls bazel-testlogs/ every 5 seconds, waits for Bazel to finish
(lock released), then uploads only new/changed results. It reads config from
.evidence/config.yaml and environment variables (EVIDENCE_STORE_URL,
EVIDENCE_STORE_API_KEY). Logs go to .evidence/watch.log; use --foreground
with watch start to run in the foreground instead, for debugging.
The commands above assume a consumer workspace with evidence_store_bazel
added as a bazel_dep; replace @evidence_store_bazel//cmd/evidence-bazel
with //cmd/evidence-bazel when running from inside adapters/bazel/ in this
repo.
Not all test workflows produce a JUnit test.xml. Failure tests (where a
bazel build is expected to fail with a specific stderr pattern) and
shell-driven integration tests determine pass/fail outside Bazel's test runner.
For these, use the record subcommand to emit a single evidence record with an
externally-determined verdict:
bazel run @evidence_store_bazel//cmd/evidence-bazel -- record \
--procedure-ref "//fire/starlark/failure_test:version_too_old_basic" \
--result PASS \
--notes "expected 'static_assert' pattern found in stderr" \
--tags failure_test,version_too_old
--evidence-type defaults to ci — a failure test is still something a machine
ran — and accepts only ci, manual_test or demonstration. It is checked
before anything is uploaded, so a wrong value fails the command instead of being
rejected record by record halfway through a loop. What kind of automated test it
was belongs in --tags, which is where it is filterable anyway.
Unlike the ingest and watch paths, record does not set metadata.collector:
a verdict decided outside Bazel's test runner was not collected by Bazel either.
Pass one yourself if it is worth recording — --metadata '{"collector":"shell"}'.
For calls in tight loops, avoid bazel run in the inner loop — its
per-invocation configuration (and any differing --action_env flags on your
other builds) churns the analysis cache. Build once up-front and exec the
binary directly:
bazel build @evidence_store_bazel//cmd/evidence-bazel
BIN="$(bazel info workspace)/$(bazel cquery --output=files @evidence_store_bazel//cmd/evidence-bazel)"
for tgt in $targets; do
"$BIN" record --procedure-ref "$tgt" --result PASS ...
done
| Flag | Default | Description |
|---|---|---|
--procedure-ref |
Target label / test identifier (required) | |
--result |
PASS, FAIL, ERROR, or SKIPPED (required, case-insensitive) |
|
--evidence-type |
ci |
How it was collected: ci, manual_test or demonstration |
--notes |
Free-text stored under metadata.notes |
|
--tags |
Comma-separated tags stored under metadata.tags |
|
--duration-ms |
Duration in milliseconds (optional) | |
--metadata |
JSON object to merge into metadata (e.g. '{"pattern":"static_assert"}') |
|
--invocation-id |
Group multiple records from the same run | |
--finished-at |
now (UTC) | RFC3339 timestamp |
--repo, --branch, --rcs-ref, --source |
auto-detected | Same as ingest path |
--api-url, --api-key |
.evidence/config.yaml |
See below |
--dry-run |
false |
Print the record as JSON instead of posting |
Config resolution order (highest priority first): command-line flag →
.evidence/config.yaml (searched upward from BUILD_WORKSPACE_DIRECTORY or
cwd). Environment variables are deliberately not consulted by record, so the
binary's behavior is stable across shell-env changes and Bazel's analysis cache
is not invalidated by invocations that happen to have different env.
INVOCATION_ID=$(uuidgen)
for tgt in $(discover_failure_tests); do
if bazel build "$tgt" 2> stderr.log; then
evidence-bazel record --procedure-ref "$tgt" --result FAIL \
--notes "expected build failure but build succeeded" \
--invocation-id "$INVOCATION_ID"
elif grep -q "static_assert" stderr.log; then
evidence-bazel record --procedure-ref "$tgt" --result PASS \
--invocation-id "$INVOCATION_ID"
else
evidence-bazel record --procedure-ref "$tgt" --result ERROR \
--notes "build failed without expected pattern" \
--invocation-id "$INVOCATION_ID"
fi
done
docker compose up -d serves the web UI from the same port as the API —
http://localhost:8000. It has four tabs: Search, Analytics, Add
Result, and Access (only shown to an administrator).
The page loads, and results can be filed, with no server to ask — which is what a test campaign at a proving ground needs. A service worker keeps the application itself (HTML, styles, scripts, icons), and a record filed with no connection goes into an outbox on the device instead of being lost.
Filing offline works like filing online: fill in Add Result and press Create. The feedback says the record was saved here rather than filed, and a counter appears in the header. From it you can see what is waiting, correct a record before it goes, or delete one.
Photos work offline too. Paste or drop an image into the test log with no connection and it is named, kept on the device, and referenced from the log straight away — the log is finished when you write it, and only the bytes are still owed. That works because a blob is named by the SHA-256 of its bytes and nothing else, so the browser can work out the reference the upload would return; nothing has to be rewritten when the record eventually goes. The photos show in the outbox, from the device's own copy, so you can see the pictures are safe and not just the words about them.
The queue is held in IndexedDB, so it survives closing the tab, quitting the browser, and restarting the machine. It also survives the login expiring, which it usually will: a session lasts 12 hours by default and a campaign does not. Sign in again and the records go.
Sending happens by itself when a connection returns — on reconnect, and on page load. Nobody is looking at the page at the moment a signal appears, so the report comes afterwards rather than as a prompt beforehand. Each record is answered on its own terms:
| What the store says | What happens |
|---|---|
| Filed | It leaves the queue |
| Already filed | It leaves the queue too — an earlier attempt got through and its response did not, which is what client_record_id is for |
| Refused (a bad field) | It stays, flagged with the store's own message, and is not retried until you change it |
| No answer at all | It stays, unchanged, and goes with the next attempt |
Photos are uploaded before the records that name them, always: a record filed first would point at bytes the store does not have, and a test log cannot be edited once filed. Uploading the same image twice costs one object, so an upload interrupted halfway through a campaign's photographs is free to repeat. Once the store has a photo and its record is filed, the device releases the bytes.
Nothing leaves the queue until the store has said what became of it, so there is no state in which the page has forgotten a record the store never received.
A record also remembers who wrote it, and is only ever sent by that person. If someone else is signed in, it waits for its author rather than being filed under the wrong name.
The weather field works offline too, by a different route. The lookup is a
server call by design, so with no connection there is nobody to ask — but the
tester is standing in the weather and does not need a model to tell them what it
is doing. Write it down opens a few boxes (conditions, temperature, wind,
humidity, rain) that compose into the same one line, in the same order and units,
that a fetched reading produces. A line written by hand carries no
weather_observed_at, which is what lets a reader tell a measurement from a
person's account of the sky.
If the field is left empty and the record names a point, the reading is fetched during the sync instead — the last moment it can be, since a filed record is immutable. Your own words are never replaced.
What does not work offline is anything that is a question about the archive: Search, Analytics, and the suggestions in the repo and procedure boxes. Those say so rather than failing obscurely. Nothing cached here is evidence — a record served from a browser cache would be a claim about the archive that might have been true last week, and a reader could not tell it from a live one.
Until a record syncs, the only copy is in this browser. The page says which of these applies rather than leaving you to find out:
| What | Effect | What the page does |
|---|---|---|
| Ordinary use | Nothing. IndexedDB is on disk and survives closing the tab, quitting the browser, and restarting the machine | Asks the browser to mark the data persistent, so eviction needs a deliberate act |
| The browser will not promise | It may reclaim the space if the device runs low | Says so in the outbox, in orange |
| No storage at all (some private windows) | The queue lasts only as long as the tab | Warns in red when the first record is queued, not later |
| iOS Safari, site not installed | Storage for a site not visited in 7 days is evicted | Add to Home Screen exempts it — see below |
| "Clear browsing data" | Gone, with everything else | Nothing can prevent it; the header counter means it is not invisible beforehand |
| A different browser, profile or device | Has its own empty outbox | The counter is per-browser |
A record that has been waiting is the failure this feature can actually produce, so the header stops counting and starts saying how long: 7 days gets a warning, 30 days turns it red. Nothing ever expires or is refused — a record queued in March is still a true account of a test that happened in March, and refusing it would destroy the only copy to punish the delay.
Two things to know before relying on it:
localhost is
exempt, which is why it works in development). The page says so in the
browser console rather than failing silently.Installed, it runs in its own window with no browser chrome, and long-pressing
the icon offers Add Result, which opens straight on the form — the only
thing anybody opens this for on a campaign. The same shortcut works as a plain
link: /#add, /#search, /#analytics.
When a newer build has been fetched, a line appears offering to reload. It only offers: reloading a page with a half-written test log in it would throw the log away, so when to take the new version is the tester's call.
The full design, including what is deliberately left out, is in docs/offline-support-plan.md.
The Add Result tab is how a person files what they ran. Repo, Commit, Procedure and Result are required; Evidence type defaults to Manual Test.
metadata.observations. Paste or drop an image into
it and it uploads, is referenced from the log, and renders in the record
dialog later; images over ~1.5 MB are downscaled in the browser first.
Images are content-addressed (named by the SHA-256 of their bytes), so the
same screenshot filed twice costs one object. Details:
Test logs and
Images in test logs.Cmd-Z/Ctrl-Z steps back through what you typed and
through the reference an upload wrote into the log. Buttons that fill a field
in for you — Now, Locate, Look up — can be undone too.The page footer names the server's build: 2026.08.27.16.23, the minute it was
made in UTC. It is what to quote when reporting that something misbehaved, and
curl http://localhost:8000/version answers the same question from a script —
see Which build is running.
Builds made with docker compose build are stamped at link time. A go run
build from a clean checkout falls back to the commit's own time, and one from a
working copy with uncommitted changes reads dev, because it matches no commit.
The footer is empty when the server has not been reached at all — a remembered
version would be a claim about a deployment nobody has checked.
The Search tab lists and filters records, with the same regex support the
API has: prefix a text filter with ~ for a POSIX regular expression instead
of an exact match, e.g. ~^release/ on branch. A single Branch, tag or
commit box matches whichever of the three the value turns out to be, instead
of asking which column to search. Selecting a row opens the record dialog, with
its test log, images, location and weather rendered alongside the record's
other fields rather than buried in a metadata dump. Full parameter list:
Querying evidence and
Regex filtering.
When an impact analysis determines that a commit's evidence is still valid for another — nothing changed in the code a test exercises — that's declared once, rather than re-run:
curl -X POST http://localhost:8000/api/v1/inheritance \
-H "Authorization: Bearer $ADMIN_KEY" \
-H 'Content-Type: application/json' \
-d '{
"repo": "myorg/firmware",
"source_rcs_ref": "abc123def",
"target_rcs_ref": "def456abc",
"justification": "Impact analysis JIRA-1234: no changes in pkg/",
"created_by": "ci-bot"
}'
From then on, every record filed against abc123def (the tested version) also
appears when querying def456abc (the version that inherits it) — in the
Search tab, checking Include inherited (on by default) shows them in a
separate panel and badges the record dialog Inherited. There is currently no
form for creating a declaration in the web UI; POST /api/v1/inheritance
above, or GET /api/v1/inheritance to list existing ones, is the only way in
and needs the admin role. Full shape and the include_inherited query
parameter: Creating an inheritance declaration.
The Analytics tab aggregates evidence instead of listing it record by record: an overview of the filtered window, a sortable per-test table (click a header to rank by fail rate, flip rate, infra errors, or reliability), and co-failure clusters with a minimal set of tests that would catch most failures. It waits for Apply before querying, since these aggregations scan far more than a search does. Selecting a row opens the matching records in Search; Export CSV downloads exactly what the table is showing. Metrics, labels and the clustering parameters are documented in full under Analytics.
The Admin tab — visible only to a caller holding principal:admin — lists
every principal with its roles, when its key was last used, and whether it's
been revoked, and is where keys are issued, rotated and revoked. It requires
EVIDENCE_AUTH_DB=true; see
Database-backed principals for the underlying
model and the equivalent /api/v1/principals calls.
bazel build //... # build everything
bazel test //... # run all tests
bazel run //cmd/server # start the server
adapters/bazel is a separate Bzlmod module (see Bazel adapter)
and is excluded from the root workspace via .bazelignore; build and test it
from inside that directory, the way CI does:
cd adapters/bazel
bazel test //...
Frontend unit tests (the web UI ships as plain ES modules, no build step) run under Node's own test runner:
node --test web/tests/*_test.mjs
prek hooks (prek.toml) run go fmt, go vet, a
build check, and basic hygiene checks (trailing whitespace, large files, merge
conflict markers) — install it and run prek install once per checkout so
these run on commit rather than in review.
docker compose up -d
./scripts/smoke-test.sh
Exercises the running API end-to-end (scripts/smoke-test.sh) — useful after
a config or migration change to check the server actually comes up serving
correctly, beyond what the unit and integration tests cover.
./scripts/dogfood.sh runs this repo's own bazel test //... and uploads the
results through the Bazel adapter — the fastest way to get evidence that looks
like a real CI run into a local store. See
Using it on your own machine for the adapter
commands it wraps.
scripts/seed-demo fills the database with synthetic evidence for demos and for
exercising the UI at realistic scale.
docker compose up -d db
go run ./scripts/seed-demo # 2,000,000 CI records + 3,000 manual tests
go run ./scripts/seed-demo --count 50000 # a smaller CI set
go run ./scripts/seed-demo --manual-tests 500 # fewer manual test logs
go run ./scripts/seed-demo --truncate # replace existing evidence
It writes to Postgres with COPY rather than through the API — the batch
endpoint inserts one row per round trip, which takes tens of minutes at this
size. API-level validation is therefore bypassed, so the generator is written to
produce records that satisfy it anyway.
Records are clustered onto a limited set of repositories, branches and commits so
that filtering returns meaningful groups, with a realistic verdict distribution
(88% PASS) and timestamps biased towards the recent past. --seed makes a run
reproducible. Two million records occupy roughly 900 MB including indexes.
Manual tests are seeded separately from that bulk CI noise: --manual-tests
(3,000 by default) generates a curated batch of manual_test records with a
fixed 50/20/20/10 pass/skip/error/fail split, markdown logs of varying length,
and synthetically rendered screenshots (never real photos, so there is no
copyright question) written through the real content-addressed blob store. It
honours the same EVIDENCE_BLOB_* variables as cmd/server — set them to match
if you want the images visible through a server pointed at S3/MinIO rather than
the local fs default.
web/static/pico.min.css is Pico CSS (MIT), vendored
rather than loaded from a CDN: the deployments that most need the offline UI
are behind a firewall or on a proving ground with no route out, and a
stylesheet that does not arrive leaves an unreadable page.
To move to a new version, fetch it and check what you got before committing it:
curl -sL "https://cdn.jsdelivr.net/npm/@picocss/pico@2.1.1/css/pico.min.css" -o web/static/pico.min.css
The current file is v2.1.1, 83,319 bytes, sha256:fbc9a63fc9fc9f72d12fd7fc9806e11fa9f77ae4f9cad146b27003a1119ba3db.
Adding any file to web/static/ also means adding it to embedsrcs in
web/BUILD.bazel and, if the page loads it, to the shell list in
web/static/sw.js. Both are checked by node --test web/tests/*_test.mjs,
which is what stops a file that works in development from going missing in a
container or on a page with no connection.
node --test web/tests/*_test.mjs covers the frontend logic that can be tested
without a browser. What it cannot cover is whether the page still works — a
missing import or a function moved into the wrong module leaves every test
passing and the page dead.
node --test web/smoke/*_test.mjs
Ten seconds against a running docker compose stack, driving headless Chrome.
It files records under smoke/browser-check; see
web/smoke/README.md for what it covers, what it leaves
behind, and why it exists.
CI runs go vet and golangci-lint, the latter pinned in
.github/workflows/ci.yml to a release built against the Go version go.mod
targets. To run the same checks locally, install that same version:
go install github.com/golangci/golangci-lint/v2/cmd/golangci-lint@v2.13.2
The pin matters more than it looks. A golangci-lint built against an older Go
than the module targets does not warn — it refuses to load the packages and
reports nothing, which is indistinguishable from a clean run if nobody reads the
output. That is how this repo went without static analysis for months.
.golangci.yml keeps the enabled set small on purpose: errcheck, govet,
ineffassign, staticcheck and unused. The stylistic ST* checks are off
because they want every package comment to begin "Package x ...", and the
comments here deliberately open by saying what the package is for.
adapters/bazel is a separate Go module and carries its own copy of the same
config, linted by its own CI job. A copy rather than a link, because the module
is published to the Bazel Central Registry on its own and has to travel with its
own settings; keep the two in step by hand.
The adapter module (evidence_store_bazel) is published to the Bazel Central Registry so consumers can pin it via bazel_dep without a git_override.
Release flow:
version in adapters/bazel/MODULE.bazel.main.evidence_store_bazel-v<version> (e.g., evidence_store_bazel-v0.0.1) and push the tag.Release evidence_store_bazel workflow (.github/workflows/release-bazel-adapter.yml) builds a source tarball of adapters/bazel/, creates a GitHub release, then calls the bazel-contrib/publish-to-bcr reusable workflow to open a PR against bazelbuild/bazel-central-registry from the fork configured in the workflow.One-time setup (before the first release):
bazelbuild/bazel-central-registry to nesono/bazel-central-registry.public_repo scope and save it as the BCR_PUBLISH_TOKEN repository secret..bcr/ at the repo root; .bcr/config.yml declares adapters/bazel as the module root.