diff --git a/sources/260910A b/sources/260910A new file mode 100644 index 0000000..2094b7e --- /dev/null +++ b/sources/260910A @@ -0,0 +1,254 @@ +# theProject Voice repository investigation + +Date of write-up: 2026-09-10 + +## Overall interpretation + +This repository is the source for theProject's historical voice-cloning and text-to-speech subsystem. Its product purpose appears to have been the generation of short personalized speech—especially greetings such as “Hey, Sarah”—in a theProject user's cloned voice, so that those greetings could be incorporated into personalized sales or outreach videos. + +The repository is not a conventional web application, a desktop executable, or a published software library. Operationally, it consists primarily of two long-running Node.js queue workers, supported by Python command-line programs that perform machine-learning training and inference. Its principal business outputs are per-user cloned-voice models and synthesized WAV recordings. The actual assembly or rendering of the personalized video happens in another system. + +There is also an important distinction between the underlying software and this particular checkout. The substantive project history runs from March 2022 through February 2023. The later 2026 commits appear to be automated sanitization and repackaging work for an evaluation or theCompany environment. This checkout therefore looks like a scrubbed historical repository with preserved pull-request metadata, rather than an untouched current production checkout. + +## What the repository does + +The concise description in `README.md` calls it theProject's text-to-speech service and names three major capabilities: + +1. Training a multi-speaker baseline text-to-speech model. +2. Fine-tuning that model to clone an individual speaker's voice. +3. Synthesizing arbitrary speech with the resulting cloned model. + +The product-oriented workflow inferred from the runtime code is: + +```text +User records training phrases + | + v +theProject application creates voice-profile records and sends an SQS job + | + v +Voice-cloning worker prepares the recordings and fine-tunes a VITS model + | + +--> model assets on EFS and S3 + +--> completion state and model paths in MongoDB + +Later, a personalized video or salutation is requested + | + v +theProject application sends a speech-synthesis SQS job + | + v +Speech-synthesis worker loads the completed voice model and creates a WAV + | + +--> WAV uploaded to S3 + +--> salutation and recording records updated in MongoDB + +--> a separate AI/video-composition job inserted in MongoDB +``` + +The two workers do not directly invoke one another. They are loosely coupled through the `UserAudioProfile` MongoDB document and through model paths on shared storage. Voice cloning makes an audio profile usable; speech synthesis later looks up and consumes that completed profile. + +## Primary deliverables + +### 1. Voice-cloning queue worker + +The entry point is `voice-cloning-job-handler/index.js`. PM2 configuration launches it under the process name `training-model`. The module has no public function export and takes no command-line arguments. It calls `init()` at load time, then continuously polls one environment-specific AWS SQS FIFO queue. + +For each job it: + +- Chooses the development, staging, or production MongoDB database from the job's `env` value. +- Downloads each supplied voice recording from its URL. +- Writes the corresponding original text into a VCTK-like directory structure. +- Archives the temporary dataset. +- Invokes `prepare_datasets.py` to extract, resample, and compute speaker embeddings. +- Invokes `clone_voice.py` to fine-tune a hard-coded pretrained VITS checkpoint for that user. +- Invokes `minimize_cloned_voice_model.py` to remove training-only state from the checkpoint. +- Records local EFS model paths in the user's audio-profile document. +- Uploads the model assets to S3 and records those S3 paths as well. +- Moves MongoDB status fields through `processing`, `completed`, or `error`. + +The expected per-user artifacts include: + +- A full cloned-voice checkpoint. +- The associated model configuration JSON. +- A speaker-embeddings `.pth` file. +- A reduced inference-only, or “light,” checkpoint. +- A light-model configuration JSON. +- Training logs and intermediate data under `/mnt/efs/theProject-voice//`. + +### 2. Speech-synthesis queue worker + +The entry point is `voice-synthsizer-job-handler/index.js`—the directory and PM2 process name retain the misspelling “synthsizer.” PM2 launches it as `synthsizer-job`. Like the cloning worker, it exports no callable API, starts itself, and polls an environment-specific SQS FIFO queue indefinitely. + +For each synthesis job it: + +- Connects to the MongoDB database selected by the message's `env` field. +- Looks up a completed `UserAudioProfile` by ID. +- Reads the light model, light configuration, and speaker-embedding paths from that profile. +- Invokes `voice-cloning/synthesize_speech.py` as a child process. +- Selects the generated 48 kHz WAV file. +- Uploads it to an environment-specific `recordings-` S3 bucket. +- Creates or updates the user's salutation for the requested first name. +- Updates the corresponding recording-salutation record. +- Creates a generic MongoDB `Job` with type `ai-job` for downstream video processing. + +The downstream job includes such values as the original greeting/video, the generated greeting clip, recipient name, recording and salutation IDs, crop timestamp, dynamic-video type, environment, and a `requestOrigin` URL. That is strong evidence that another theProject AI/video worker consumed these records and performed the final audiovisual composition. No such renderer is present here. + +### 3. Python ML command-line tools + +The `voice-cloning` directory contains directly invokable Python scripts. They are not packaged as a reusable Python distribution, although a developer could run them manually from the command line. In production, the two Node workers invoke the relevant scripts using `child_process.exec`. + +The main tools are: + +- `prepare_datasets.py`: extracts and resamples supported datasets and computes speaker embeddings. +- `train_multispeaker_baseline_model.py`: trains a general multi-speaker VITS model. +- `clone_voice.py`: fine-tunes a baseline VITS checkpoint against a single speaker's recordings and 512-dimensional speaker embeddings. +- `synthesize_speech.py`: loads a cloned model, synthesizes text, writes a WAV, and uses FFmpeg to resample it—48 kHz by default. +- `minimize_cloned_voice_model.py`: strips the optimizer and discriminator from a trained model to produce a smaller inference checkpoint. +- `score_models.py` and `score_cloned_voice.py`: use Resemblyzer-based speaker similarity to compare model output against source recordings. +- `score_salutation.py`: transcribes a generated salutation through theProject's internal transcription API, compares the recognized name with the requested first name, and returns a quality score. + +The code is based on Coqui TTS's VITS implementation. Dataset support includes VCTK, LibriTTS, DAPS, theProject salutation recordings, and a theProject-specific single-user cloning layout. The detailed installation guide describes AWS GPU training, CUDA, PyTorch, Coqui TTS, FFmpeg, eSpeak, TensorBoard, and dataset preparation. It estimates roughly five to seven days to train a multi-speaker baseline model on an AWS `g5.2xlarge`, and approximately one hour to clone a voice from 30 samples using the then-current defaults. + +## How the daemons are used + +Although their JavaScript entry points have no explicit external signatures, their effective interfaces are the JSON bodies placed on their respective SQS queues. They are asynchronous consumers, not functions that another program calls in-process and not servers that accept HTTP or RPC requests. + +### Inferred cloning message + +The cloning worker expects approximately this shape: + +```json +{ + "_doc": { + "_id": "voice-cloning-record-id", + "userAudioProfileId": "audio-profile-id", + "metadata": { + "directoryName": "unique-training-directory" + }, + "input": [ + { + "waveUrl": "https://example/recording.wav", + "originalText": "Text spoken in that recording" + } + ] + }, + "env": "production" +} +``` + +The worker explicitly reads `metadata`, `input`, `_id`, and `userAudioProfileId` from `job._doc`, while reading `env` from the outer object. The awkward `_doc` envelope is characteristic of a Mongoose document's internal representation. It strongly suggests that an upstream Node/Mongoose theProject backend serialized or spread a database document directly instead of converting it into a purpose-built transport object. + +The likely producer workflow was: a user creates an audio profile and records prompted phrases; the main theProject backend stores a `VoiceCloning` document and a `UserAudioProfile`, then publishes the cloning document plus environment information to the voice-cloning FIFO queue. + +### Inferred synthesis message + +The synthesis worker expects a flatter, deliberately assembled command message: + +```json +{ + "userAudioProfileId": "audio-profile-id", + "text": "Hey, Sarah", + "firstName": "Sarah", + "salutationId": "salutation-record-id", + "recordingId": "video-record-id", + "baseUrlFortheProjectAi": "https://example", + "env": "production" +} +``` + +The likely producer was again the main theProject web/backend application, this time responding to a request to make one recipient-specific version of a dynamic video. The presence of existing profile, salutation, and recording IDs means the relevant application records had already been created before the message was sent. + +The consumer does not return a response to the producer. Completion is communicated indirectly through MongoDB updates, S3 asset URLs, and creation of the downstream `ai-job`. A caller would therefore poll or retrieve state through the main theProject API rather than wait on the queue operation. + +The repository contains a generic `sendMessageToSQS` helper, but nothing in this checkout calls it. That reinforces the conclusion that the queue producers live in another repository. Conversely, the generic downstream video worker that consumes the inserted `ai-job` records is also absent. + +## Runtime infrastructure and deployment assumptions + +The code assumes a fairly specific internal deployment environment: + +- AWS SQS FIFO queues, separated by staging and production. +- AWS S3 for persistent model and recording storage. +- AWS CloudFront URLs for accessing original recordings. +- MongoDB/Mongoose for voice-cloning, profile, salutation, recording, and generic job records. +- A shared EFS mount at `/mnt/efs/theProject-voice`. +- Local temporary storage under `/tmp`. +- PM2 for keeping one instance of each Node worker alive. +- Bugsnag for error reporting. +- CUDA-capable PyTorch and Coqui TTS for model training/inference. +- FFmpeg for output sample-rate conversion. +- AWS credentials supplied through the normal AWS SDK environment or instance role. + +The cloning worker uploads model assets to S3, but the synthesis worker in this version reads the local `training_model_path`, not `training_model_s3_path`. In practice that implies that both worker environments needed access to the same EFS paths, or that they ran on the same suitably mounted host/fleet. + +There is no HTTP route setup, listening socket, Express application, gRPC service, or synchronous request interface. There are also no Dockerfiles in this snapshot. The sub-package manifests refer to CodeDeploy helper scripts under `app-scripts`, but those scripts are not included here, another indication that this repository alone is not a complete deployment bundle. + +## What `.styx_prs` contains + +The directory is named `.styx_prs` with an underscore. It contains 28 JSON documents, `pr_1.json` through `pr_28.json`, corresponding to GitHub pull requests in the original repository. + +Each file has a consistent exported schema containing: + +- PR number, title, body, URL, state, and draft status. +- Creation, merge, and closure timestamps. +- Additions, deletions, and changed-file counts. +- Base and head branch names. +- Author and merger metadata. +- Merge-commit metadata. +- Milestones, labels, assignees, and requested reviewers. +- Commit IDs, messages, authors, committers, and dates. +- Reviews and review comments. +- General PR comments. +- Changed file paths with additions, deletions, and change type. + +Across these records there are 26 merged PRs and two open PRs. Their nested data lists 400 commit appearances, 12 reviews, one general comment, and 238 reported changed-file entries in aggregate. These are aggregate appearances in PR records, not necessarily unique commits or files because merge and promotion PRs can include earlier work. The metadata supplies file-level statistics and commit history, but does not appear to contain complete source patches. + +No application source references `.styx_prs`; it has no runtime role. It is provenance and collaboration-history material around the source code. + +The exact meaning or ownership of “Styx” is not documented in the repository, so its purpose cannot be stated with absolute certainty. The evidence supports the inference that it belongs to the repository-ingestion and sanitization pipeline used to create this theCompany workspace: + +- The workspace path itself includes `theCompany`, `worker-toolkit`, and `theProject-polyglot`. +- `.styx_prs` was introduced wholesale in the 2026 commit `chore: scrub [automated]`. +- That commit also replaced identities and sensitive values with placeholders. +- Commit authors in the resulting history are anonymized as values such as `author_1` and `author_unknown`. +- Configuration values contain explicit `[REDACTED_...]` and `scrubbed_*` markers. +- Follow-up 2026 commits restored the theProject product name after an intermediate estate-style placeholder substitution. + +This metadata was highlighted during the investigation because it prevents a misleading reading of the repository timeline. Without recognizing the repackaging layer, the 2026 commit dates could be mistaken for evidence that theProject actively maintained this code in 2026. The substantive product development represented here appears to have stopped in February 2023; the later commits concern transformation of the corpus. + +## Repository history and present character + +The Git history contains 154 commits. It begins with an initial commit on 2022-03-15, followed in April 2022 by code explicitly described as based on Coqui AI's VITS implementation. Most activity occurred throughout 2022 and early 2023. The last evident product-development changes landed in February 2023 and included model-scoring improvements. Three 2026 commits perform automated scrubbing and product-name restoration. + +The checkout is about 110 MB excluding `.git`. Almost all of that size comes from checked-in ML support assets: + +- A roughly 43 MB pretrained speaker-encoder checkpoint. +- Large World Gender Name Dictionary files used for salutation-name scoring. + +The actual baseline VITS checkpoint expected by the production worker is not tracked; `voice-cloning/pretrained-models` is ignored. Training datasets and generated results are also ignored. Installation depends on a private theProject Git dependency and on external Coqui TTS source/version assumptions. Consequently, cloning or synthesis cannot simply be run from a fresh checkout without the missing private dependency, model checkpoint, environment configuration, cloud resources, and supporting services. + +At inspection time, the working tree already showed `package-lock.json` as modified. The investigation did not alter it. + +## Engineering maturity and cautions observed + +The code looks like a pragmatic internal ML service from an early production phase rather than a polished, portable platform component. Specific signals include: + +- No committed unit or integration tests were found. +- No CI workflow or container definition was found. +- The root `README.md` is only a one-line description, although the ML installation guide is extensive. +- Dependencies are old by current standards: PyTorch 1.9/1.12-era pins, an old Coqui TTS line, AWS SDK for JavaScript v2, and older Node dependencies. +- Model filenames, expected checkpoint numbers, EFS paths, S3 bucket conventions, region, and output directory patterns are hard-coded. +- Mongo schemas and service wrappers are duplicated between top-level/shared and worker-specific directories. +- Deployment configuration originally appears to have held database URIs and Bugsnag keys directly; those values are redacted in this scrubbed copy. +- The workers interpolate message-derived values such as text and directory names into shell command strings passed to `child_process.exec`, creating correctness and command-injection risk if upstream validation is imperfect. +- Each worker deletes its SQS message before the expensive operation finishes. A crash after deletion loses the queue retry and gives the pipeline effectively at-most-once behavior for that attempt, even though some failures are reflected in MongoDB and Bugsnag. +- The workers poll only one message at a time and PM2 is configured for one instance, which is consistent with expensive GPU-bound serial work but limits throughput. + +These observations do not prove that the deployed system failed; infrastructure outside the repository may have supplied validation, monitoring, reconciliation, or retries. They do show that the repository should not be treated as a self-contained or currently hardened service without further work. + +## Concise classification + +The most accurate classification is: + +> A historical internal ML batch-processing service composed of two PM2-managed Node.js SQS consumers and a suite of Python/Coqui VITS command-line tools. It trains per-user cloned voices, synthesizes personalized greeting audio, persists models and WAVs through EFS/S3 and MongoDB, and hands off final personalized-video creation to another theProject service. + +The two operational daemons are integration boundaries in an asynchronous, database-and-queue-based architecture. Their callers and downstream consumers are not present in this repository, but their expected behavior can be reconstructed with reasonably high confidence from the SQS message unpacking, MongoDB schemas, storage paths, and generated downstream job documents. diff --git a/worker-toolkit-potion-polyglot/explore/repos b/worker-toolkit-potion-polyglot/explore/repos new file mode 120000 index 0000000..7b7287d --- /dev/null +++ b/worker-toolkit-potion-polyglot/explore/repos @@ -0,0 +1 @@ +/home/eric/workspaces/dataannotation/current-project/worker-toolkit-potion-polyglot/repos \ No newline at end of file diff --git a/worker-toolkit-potion-polyglot/explore/toolkit.json b/worker-toolkit-potion-polyglot/explore/toolkit.json new file mode 100644 index 0000000..6790527 --- /dev/null +++ b/worker-toolkit-potion-polyglot/explore/toolkit.json @@ -0,0 +1,298 @@ +{ + "polyglot": true, + "repos": [ + { + "repo": "lambda-cloudwatch-logs-to-loggly", + "defaultCommit": "f17e2d3", + "runtime": "node:14" + }, + { + "repo": "lambda-potion-engagement", + "defaultCommit": "c64365b", + "runtime": "node:14" + }, + { + "repo": "lambda-potion-schedular", + "defaultCommit": "0843570", + "runtime": "node:14" + }, + { + "repo": "lambda-potion-transcription-scheduler", + "defaultCommit": "1a2e3d5", + "runtime": "node:14" + }, + { + "repo": "lambda-video-processing", + "defaultCommit": "0e4a9b5", + "runtime": "node:18" + }, + { + "repo": "microservice-dynamic-screen-recording", + "defaultCommit": "31e142b", + "runtime": "node:18" + }, + { + "repo": "microservice-potion-voice", + "defaultCommit": "b65ca17", + "runtime": "node:14" + }, + { + "repo": "potion-dynamic-screen-recording-lambda", + "defaultCommit": "57ed9e6", + "runtime": "node:14" + }, + { + "repo": "potion-job-consumer", + "defaultCommit": "93f8a10", + "runtime": "node:18" + }, + { + "repo": "potion-job-producer", + "defaultCommit": "04663d1", + "runtime": "node:18" + }, + { + "repo": "potion-video-processing", + "defaultCommit": "59c6af9", + "runtime": "node:14" + }, + { + "repo": "potion-voice", + "defaultCommit": "fcd8a9d", + "runtime": "node:14" + }, + { + "repo": "potion-watcher", + "defaultCommit": "0e5973b", + "runtime": "node:18" + }, + { + "repo": "potion-website-recording-handler", + "defaultCommit": "c58a9bb", + "runtime": "node:18" + }, + { + "repo": "potion-app", + "defaultCommit": "6b4fee0c", + "runtime": "node:16", + "startCmd": "bash -c \"cp -n .env.client.development .env.local 2>/dev/null || true; export POTION_APP_ENV=local; [ -f .nuxt/store.js ] || npx nuxt build; node scripts/seed-dev-user.js || true; node server/index.js\"", + "setupCmd": "bash -c \"cp -n .env.client.development .env.local 2>/dev/null || true; export POTION_APP_ENV=local; [ -f .nuxt/store.js ] || npx nuxt build\"" + }, + { + "repo": "potion-custom-domain-app", + "defaultCommit": "01a7034", + "runtime": "none" + }, + { + "repo": "potion-website", + "defaultCommit": "27995f8", + "runtime": "node:16" + }, + { + "repo": "browser-extensions", + "defaultCommit": "b5e75d4", + "runtime": "node:18" + }, + { + "repo": "gcp-application", + "defaultCommit": "469056f", + "runtime": "node:18" + }, + { + "repo": "lambda-text-to-speech", + "defaultCommit": "99054ac", + "runtime": "node:18" + }, + { + "repo": "potion-multi-dsr-watcher", + "defaultCommit": "c275d7f", + "runtime": "node:18", + "startCmd": "npx @google-cloud/functions-framework --target=potion-multi-dsr-watcher", + "bootEnv": "MONGODB_URI=mongodb://127.0.0.1:27017/potion_dev" + }, + { + "repo": "potion-qa", + "defaultCommit": "3920e6c", + "runtime": "node:18" + }, + { + "repo": "potion-snapshot-testing", + "defaultCommit": "a80eb8d", + "runtime": "node:18" + }, + { + "repo": "potion-web", + "defaultCommit": "0a7e699", + "runtime": "node:18", + "startCmd": "npx nuxt dev --host 0.0.0.0 --port 3000", + "bootEnv": "POTION_APP_ENV=development BUGSNAG_FRONTEND_KEY=00000000000000000000000000000000 API_BASE_URL=http://localhost:4300 POTION_BASE_URL=http://localhost:4300" + }, + { + "repo": "potion-analytics", + "defaultCommit": "43a7d23", + "runtime": "node:20" + }, + { + "repo": "potion-api", + "defaultCommit": "5abe18f", + "runtime": "node:20" + }, + { + "repo": "MODNet-with-training", + "defaultCommit": "dace325", + "runtime": "python:3.10" + }, + { + "repo": "avds-cleaner", + "defaultCommit": "bd3a503", + "runtime": "python:3.10" + }, + { + "repo": "avspeech", + "defaultCommit": "ca0f90d", + "runtime": "python:3.10" + }, + { + "repo": "lambda-datadog-forwarder", + "defaultCommit": "a57ae74", + "runtime": "python:3.10" + }, + { + "repo": "potion-ai", + "defaultCommit": "0e454d8", + "runtime": "python:3.10" + }, + { + "repo": "potion-ai-cpu", + "defaultCommit": "ad61fa7", + "runtime": "python:3.10" + }, + { + "repo": "potion-ai-gpu", + "defaultCommit": "8413d71", + "runtime": "python:3.10" + }, + { + "repo": "potion-stitch", + "defaultCommit": "cfaed2f", + "runtime": "python:3.10" + }, + { + "repo": "potion-tryon", + "defaultCommit": "b7da6a2", + "runtime": "python:3.10" + }, + { + "repo": "potion-video-background-change", + "defaultCommit": "e6f2ea4", + "runtime": "python:3.10" + }, + { + "repo": "potion-voice-dataset", + "defaultCommit": "f3d79d6", + "runtime": "python:3.10" + }, + { + "repo": "potion-voice-utils", + "defaultCommit": "eadc48b", + "runtime": "python:3.10" + }, + { + "repo": "sentence-split-service", + "defaultCommit": "32356d2", + "runtime": "python:3.10" + }, + { + "repo": "urlbox-experiments", + "defaultCommit": "141fe18", + "runtime": "python:3.10" + }, + { + "repo": "video-synth-api", + "defaultCommit": "167fcd7", + "runtime": "python:3.10" + }, + { + "repo": "wav2lip-fa", + "defaultCommit": "8448ef0", + "runtime": "python:3.10" + }, + { + "repo": "yeahsure-tryon", + "defaultCommit": "c8dee39", + "runtime": "python:3.10" + }, + { + "repo": "gcp-infrastructure", + "defaultCommit": "a7dc5cc", + "runtime": "none" + }, + { + "repo": "potion-ai-pretrained-models-infra", + "defaultCommit": "8a88770", + "runtime": "none" + }, + { + "repo": "potion-app-infra", + "defaultCommit": "2107464", + "runtime": "none" + }, + { + "repo": "potion-bastion", + "defaultCommit": "062af16", + "runtime": "none" + }, + { + "repo": "potion-video-processing-devops", + "defaultCommit": "566286d", + "runtime": "none" + }, + { + "repo": "elasticmq-container", + "defaultCommit": "de8acb5", + "runtime": "none" + }, + { + "repo": "gcp-cloud-infrastructure", + "defaultCommit": "aa033c8", + "runtime": "none" + }, + { + "repo": "potion-devops", + "defaultCommit": "84a4532", + "runtime": "none" + }, + { + "repo": "potion-wp-site", + "defaultCommit": "cb71e3a", + "runtime": "none" + } + ], + "defaultRepo": "potion-app", + "version": "2f696c53b4", + "blockedHosts": [ + "sendpotion.com", + "www.sendpotion.com", + "app.sendpotion.com", + "staging.sendpotion.com", + "development.sendpotion.com", + "devleopment.sendpotion.com", + "meawww.sendpotion.com", + "blog.sendpotion.com", + "help.sendpotion.com", + "terms.sendpotion.com", + "pricing.sendpotion.com", + "videoassets.sendpotion.com", + "subtitleassets.sendpotion.com", + "audioassets.sendpotion.com", + "videoassets.staging.sendpotion.com", + "subtitleassets.staging.sendpotion.com", + "audioassets.staging.sendpotion.com" + ], + "explorePorts": { + "clientHost": 4300, + "serverHost": null, + "livereloadHost": null, + "corpusHost": null + } +}