Corpus viewer + search index
Some toolkits ship a reference-data corpus — the source company's real internal data
(chat exports, tickets, email, support conversations, docs) at data/zeta-corpus/ (mounted
at /data/zeta-corpus in the containers). This directory adds two ways to digest it:
- A web viewer — full-text search across every source at once, browsable channels /
projects / mailboxes, per-person activity pages, cross-references (a ticket's slack
chatter, a message's ticket), and day views ("everything that happened on 2023-02-09" —
pairs well with a commit date from
git log). - A queryable index —
data/corpus-index/corpus.db, one SQLite database (FTS5) with every message / issue / comment / email / doc normalized into a singledocstable. Point Claude (orsqlite3, or python) at it directly instead of grepping 130k files.
The index is built from the exact corpus this toolkit ships — it adds no data that
isn't already under data/zeta-corpus/, it just makes it findable. Doc rows carry a
meta.path (or a derivable path — see below) back to the raw file, and the viewer shows
it on every doc, so you can jump from a search hit to the underlying export when you
write a task prompt around it.
The viewer
In the Explore container it starts automatically; view-corpus prints the URL
(view-corpus --help for start/stop/logs). Anywhere else:
python3 explore/corpus-viewer/serve.py # from the toolkit root; zero dependencies
and open the printed URL. If this toolkit build has no data/corpus-index/, the viewer
has nothing to serve (rebuilt toolkit downloads include it).
The index schema (corpus.db)
One row per atomic item in docs:
| column | meaning |
|---|---|
source |
slack | jira | github | custom_chat | gmail | takeout | confluence |
kind |
message/bot/system (slack), issue/comment (jira), pr/review/review_comment/commit (github), conversation/chat_message/email/complaint (support), email (gmail), file/page (docs) |
container |
slack channel / jira project / repo / mailbox / drive account / wiki space (support for custom_chat) |
ext_id |
stable id: jira key (ENG-3051), slack ts, message-id, repo#123, commit sha |
parent_id |
thread parent / issue / PR / conversation this doc belongs to (docs.id) |
author |
pseudonymized entity (EMPLOYEE_0011) — the same id is the same person in every source |
ts |
UTC YYYY-MM-DDTHH:MM:SS (lexically sortable) |
title |
subject / issue summary / PR title (null for chat messages) |
text |
normalized plain text |
meta |
JSON: status, priority, labels, state, attachments, path (raw file under the corpus root), … |
Support tables: docs_fts (FTS5 over title+text), containers (per-channel/project doc
counts + date ranges), entities (per-person authored/mention counts + github profile
where known), xrefs(doc_id, kind, key) (extracted references: jira keys, pr numbers,
commit sha prefixes, entity mentions), meta (build info).
meta.path is omitted where it's derivable: slack → slack/<container>/<YYYY-MM-DD>.json
(date from ts), custom_chat → custom_chat/<conversations|comments|cx_emails|cx_complaints>.csv.
Query it directly (great for Claude)
sqlite3 /workspace/data/corpus-index/corpus.db # or python3 -c "import sqlite3; …"
-- full-text search, best first (bm25): what broke around idempotency?
SELECT d.source, d.container, d.ts, snippet(docs_fts, 1, '[', ']', '…', 12)
FROM docs_fts JOIN docs d ON d.id = docs_fts.rowid
WHERE docs_fts MATCH 'idempotency NEAR key' ORDER BY bm25(docs_fts) LIMIT 20;
-- everything that references a jira ticket (slack chatter, other tickets)
SELECT d.source, d.container, d.ts, substr(d.text, 1, 120)
FROM xrefs x JOIN docs d ON d.id = x.doc_id
WHERE x.kind = 'jira' AND x.key = 'ENG-3051';
-- a person's activity across every system
SELECT source, kind, COUNT(*) FROM docs WHERE author = 'EMPLOYEE_0011' GROUP BY 1, 2;
-- what happened the week of a commit you're eyeing
SELECT source, container, COUNT(*) FROM docs
WHERE ts BETWEEN '2023-03-20' AND '2023-03-27' GROUP BY 1, 2 ORDER BY 3 DESC;
-- a slack thread, in order (parent id from the parent's docs.id)
SELECT ts, author, text FROM docs WHERE id = :parent OR parent_id = :parent ORDER BY ts;
FTS5 syntax works in MATCH: "exact phrase", term1 AND term2, NEAR(a b, 5), wire*.
Why this matters for task authoring
The corpus is what was really happening around the code you have in repos/ — incidents
in errors-production, design debates in engineering, the support fallout of real bugs
in custom_chat, the tickets that tracked them in jira. A task grounded in one of those
moments ("this complaint came in; find what shipped that week and assess the fix") is far
richer than one invented from the diff alone. Use the viewer to find the moment; use the
raw /data/zeta-corpus/... paths in your task materials.