Files
project-work/worker-toolkit-stocks-in-the-future/explore/corpus-viewer

Corpus viewer + search index

Some toolkits ship a reference-data corpus — the source company's real internal data (chat exports, tickets, email, support conversations, docs) at data/zeta-corpus/ (mounted at /data/zeta-corpus in the containers). This directory adds two ways to digest it:

  1. A web viewer — full-text search across every source at once, browsable channels / projects / mailboxes, per-person activity pages, cross-references (a ticket's slack chatter, a message's ticket), and day views ("everything that happened on 2023-02-09" — pairs well with a commit date from git log).
  2. A queryable index — data/corpus-index/corpus.db, one SQLite database (FTS5) with every message / issue / comment / email / doc normalized into a single docs table. Point Claude (or sqlite3, or python) at it directly instead of grepping 130k files.

The index is built from the exact corpus this toolkit ships — it adds no data that isn't already under data/zeta-corpus/, it just makes it findable. Doc rows carry a meta.path (or a derivable path — see below) back to the raw file, and the viewer shows it on every doc, so you can jump from a search hit to the underlying export when you write a task prompt around it.

The viewer

In the Explore container it starts automatically; view-corpus prints the URL (view-corpus --help for start/stop/logs). Anywhere else:

python3 explore/corpus-viewer/serve.py        # from the toolkit root; zero dependencies

and open the printed URL. If this toolkit build has no data/corpus-index/, the viewer has nothing to serve (rebuilt toolkit downloads include it).

The index schema (corpus.db)

One row per atomic item in docs:

column meaning
source slack | jira | github | custom_chat | gmail | takeout | confluence
kind message/bot/system (slack), issue/comment (jira), pr/review/review_comment/commit (github), conversation/chat_message/email/complaint (support), email (gmail), file/page (docs)
container slack channel / jira project / repo / mailbox / drive account / wiki space (support for custom_chat)
ext_id stable id: jira key (ENG-3051), slack ts, message-id, repo#123, commit sha
parent_id thread parent / issue / PR / conversation this doc belongs to (docs.id)
author pseudonymized entity (EMPLOYEE_0011) — the same id is the same person in every source
ts UTC YYYY-MM-DDTHH:MM:SS (lexically sortable)
title subject / issue summary / PR title (null for chat messages)
text normalized plain text
meta JSON: status, priority, labels, state, attachments, path (raw file under the corpus root), …

Support tables: docs_fts (FTS5 over title+text), containers (per-channel/project doc counts + date ranges), entities (per-person authored/mention counts + github profile where known), xrefs(doc_id, kind, key) (extracted references: jira keys, pr numbers, commit sha prefixes, entity mentions), meta (build info).

meta.path is omitted where it's derivable: slack → slack/<container>/<YYYY-MM-DD>.json (date from ts), custom_chat → custom_chat/<conversations|comments|cx_emails|cx_complaints>.csv.

Query it directly (great for Claude)

sqlite3 /workspace/data/corpus-index/corpus.db   # or python3 -c "import sqlite3; …"
-- full-text search, best first (bm25): what broke around idempotency?
SELECT d.source, d.container, d.ts, snippet(docs_fts, 1, '[', ']', '…', 12)
FROM docs_fts JOIN docs d ON d.id = docs_fts.rowid
WHERE docs_fts MATCH 'idempotency NEAR key' ORDER BY bm25(docs_fts) LIMIT 20;

-- everything that references a jira ticket (slack chatter, other tickets)
SELECT d.source, d.container, d.ts, substr(d.text, 1, 120)
FROM xrefs x JOIN docs d ON d.id = x.doc_id
WHERE x.kind = 'jira' AND x.key = 'ENG-3051';

-- a person's activity across every system
SELECT source, kind, COUNT(*) FROM docs WHERE author = 'EMPLOYEE_0011' GROUP BY 1, 2;

-- what happened the week of a commit you're eyeing
SELECT source, container, COUNT(*) FROM docs
WHERE ts BETWEEN '2023-03-20' AND '2023-03-27' GROUP BY 1, 2 ORDER BY 3 DESC;

-- a slack thread, in order (parent id from the parent's docs.id)
SELECT ts, author, text FROM docs WHERE id = :parent OR parent_id = :parent ORDER BY ts;

FTS5 syntax works in MATCH: "exact phrase", term1 AND term2, NEAR(a b, 5), wire*.

Why this matters for task authoring

The corpus is what was really happening around the code you have in repos/ — incidents in errors-production, design debates in engineering, the support fallout of real bugs in custom_chat, the tickets that tracked them in jira. A task grounded in one of those moments ("this complaint came in; find what shipped that week and assess the fix") is far richer than one invented from the diff alone. Use the viewer to find the moment; use the raw /data/zeta-corpus/... paths in your task materials.