# Potion Voice — Overview > An asynchronous voice-cloning and text-to-speech service for Potion's personalized-video pipeline, combining Node.js queue workers with a GPU-oriented Coqui VITS training and inference toolkit. ## Purpose Potion Voice has no HTTP server or user interface. It provides two continuously running workers: one fine-tunes a per-user voice model from uploaded recordings, and one uses that model to synthesize a personalized greeting and enqueue downstream video-compositing work. The repository also contains Python command-line tools for preparing speech datasets, training the shared multi-speaker baseline, cloning and minimizing individual voices, synthesizing speech, and scoring model or salutation quality. ## Tech Stack | Layer | Technology | | --- | --- | | Worker runtime | Node.js, CommonJS modules; no Node version is declared | | Process management | PM2, one process per worker | | ML runtime | Python 3 (the guide targets 3.10), PyTorch, Coqui TTS/Trainer | | Speech model | VITS with 512-dimensional speaker d-vectors; 22,050 Hz training/inference output | | Audio processing | Coqui resampling/embedding tools, `ffmpeg` for 48 kHz output, `espeak-ng` as the documented phoneme backend | | Database | MongoDB through Mongoose 6.x | | Queue and object storage | AWS SDK v2, SQS, S3, CloudFront-hosted source audio | | Compute and filesystem | GPU-backed EC2 is the documented target; trained assets and logs are placed on an EFS mount | | Monitoring | Bugsnag for worker exceptions; TensorBoard/TensorBoardX for training runs | | Evaluation | Resemblyzer speaker similarity, `textdistance`, and Potion's internal transcription API | | Tests | No automated test framework, test files, lint command, or CI configuration is present | Python dependency sets are split across `requirements*.txt`: development pins PyTorch 1.12.1/CUDA 11.6, the legacy/default set pins PyTorch 1.9.1/CUDA 11.1, production has separate CPU and unpinned-GPU variants, and local development leaves PyTorch unpinned. Every set also installs a private `potion-voice-utils` Git dependency, although this checkout has no direct import from it. ## Directory Structure ```text . ├── app/services/ Shared Node.js helpers │ ├── s3/ S3 upload/download wrapper │ ├── sqs/ SQS receive/delete/send wrapper │ ├── utils/ Error serialization, Bugsnag helper, file deletion │ └── voice_cloning/ Older duplicate VoiceCloning model/service ├── voice-cloning-job-handler/ Per-user model-training worker │ ├── index.js Queue loop and end-to-end orchestration │ ├── user_audio_profile/ Mongoose schema and CRUD service │ ├── voice_cloning/ Mongoose schema and CRUD service │ └── pm2-{development,production}.yml ├── voice-synthsizer-job-handler/ Greeting-synthesis worker (directory typo is historical) │ ├── index.js Queue loop, synthesis, upload, downstream job creation │ ├── job/ Downstream AI job schema/service │ ├── recording/ Large shared Recording schema │ ├── recording_salutation/ Dynamic-video salutation schema │ ├── salutation/ Reusable generated-salutation schema/service │ ├── user_audio_profile/ Duplicate profile schema/service │ └── pm2-{development,production}.yml ├── voice-cloning/ Python ML and audio toolkit │ ├── assets/ Speaker encoder and World Gender Name Dictionary data │ ├── docs/ EC2 setup and command examples │ ├── utils/ Synthesis, similarity, name matching, transcription helpers │ ├── prepare_datasets.py Archive extraction, resampling, d-vector generation │ ├── train_multispeaker_baseline_model.py │ ├── clone_voice.py Fine-tunes the baseline for one speaker │ ├── minimize_cloned_voice_model.py Removes training-only model state │ ├── synthesize_speech.py Generates and resamples a WAV │ └── score_*.py Manual model/salutation evaluation tools ├── requirements*.txt Python environment variants ├── package.json Shared/root Node dependencies └── README.md One-line project description ``` This is not configured as an npm workspace. There are three package manifests with largely duplicated dependencies; the worker code resolves shared modules and, depending on installation layout, dependencies from the repository root. ## Architecture ### Queue contracts | Worker | Expected SQS message body | | --- | --- | | Voice cloning | JSON with `job._doc._id`, `job._doc.userAudioProfileId`, `job._doc.metadata.directoryName`, `job._doc.input[]`, and top-level `job.env`. Each input item contains `waveUrl` and `originalText`. | | Synthesis | JSON with `userAudioProfileId`, `text`, `firstName`, `salutationId`, `recordingId`, `baseUrlForPotionAi`, and `env`. | In both workers, the message's `env` selects the Mongo URI and environment-specific storage resources. This is separate from the process-level environment used to configure PM2 and Bugsnag. ### Voice-cloning flow 1. `voice-cloning-job-handler/index.js` short-polls one message from the configured SQS FIFO queue and immediately deletes it. 2. It selects a MongoDB connection and CloudFront base URL from the message environment, then marks both the `VoiceCloning` and `UserAudioProfile` documents as `processing`. 3. It rewrites each recording URL's host to the selected CloudFront host, downloads WAV files over HTTPS, and writes a VCTK-style dataset under `/tmp//{wav48,txt}/1/`. Files are numbered `1_001`, `1_002`, and so on. 4. It archives the dataset and invokes three Python programs as child processes: - `prepare_datasets.py` computes speaker embeddings at 16 kHz, then restores and resamples the training audio to 22,050 Hz. - `clone_voice.py` fine-tunes the hard-coded `pretrained-models/checkpoint_365000.pth` VITS baseline. Defaults are batch size 96, 200 epochs, mixed precision, two evaluation samples, and checkpoints every 200 steps. - `minimize_cloned_voice_model.py` reloads `checkpoint_365200.pth`, drops the discriminator and optimizer state, and creates `_light.pth` plus `config_light.json` inference assets. 5. Generated datasets, checkpoints, configs, embeddings, and command logs live under `/mnt/efs/potion-voice///`. Mongo status moves to `completed`, and `UserAudioProfile.training_model_path` records five local paths (full/light model, full/light config, and speaker embeddings). 6. The same five files are uploaded through S3 and their returned locations are stored in `training_model_s3_path`. The code constructs the bucket argument as `potion-voice-users-training-model/` and object keys as `/`. An exception after Mongo connects marks both records `error` and reports to Bugsnag. There is no compensating queue retry because receipt deletion happens before processing. ### Greeting-synthesis flow 1. `voice-synthsizer-job-handler/index.js` receives and immediately deletes one SQS message, connects to the Mongo database selected by `job.env`, and finds a completed `UserAudioProfile`. 2. It reads the **local EFS paths** from `training_model_path`; `training_model_s3_path` is not used for inference. `synthesize_speech.py` loads the light VITS model and the profile's single-speaker embeddings, writes a native-rate WAV, and runs `ffmpeg` to create the default 48,000 Hz WAV. 3. The resampled file is uploaded to bucket `recordings-` with a generated key ending in `_salutation_.wav`. 4. The worker upserts a reusable `Salutations` record keyed by user, audio profile, and first name; updates the requested `recording_salutations` record; and loads the associated `Recordings` document. 5. It inserts a new `Job` (default type `ai-job`) containing the original video/greeting, crop timestamp, synthesized greeting URL, request origin, environment, recording IDs, and dynamic-video type. Another service is expected to consume this Mongo-backed job and composite the final personalized video. Both workers run serially in an infinite loop. Empty polls sleep for two seconds; active queues are processed without that delay. They open and close Mongoose around each message rather than maintaining a process-wide connection. ### Python toolkit The Python scripts are also usable independently from `voice-cloning/`: - Baseline training combines VCTK 0.92, LibriTTS train-clean-360, and Potion salutation recordings into a multi-speaker VITS model. The checked-in configuration targets 22,050 Hz audio and 512-dimensional d-vectors. The guide estimates 5–7 days for 100 epochs on an AWS `g5.2xlarge`. - Per-user cloning expects matching transcripts and recordings in `txt/1/` and `wav48/1/`; the guide recommends 30 samples and says a default clone takes about one hour on `g5.2xlarge`. - `score_cloned_voice.py` and `score_models.py` synthesize fixed sentences and compare Resemblyzer embeddings against real recordings; the latter ranks checkpoint files and reports a top five. - `score_salutation.py` transcribes a WAV, extracts candidate names, validates them against the included World Gender Name Dictionary, and combines transcription confidence with Jaro-Winkler, Levenshtein, and Match Rating Approach similarity. ## Integrations | Integration | Use and code location | | --- | --- | | AWS SQS (`us-west-2`) | Environment-specific FIFO queues feed both workers. Shared wrappers are in `app/services/sqs/`; queue URLs are supplied by PM2 configuration. | | AWS S3 | `app/services/s3/index.js` uploads trained model assets and synthesized greetings. AWS credentials are not explicit variables; the AWS SDK's normal credential chain is assumed. | | CloudFront/HTTPS | The cloning worker replaces the host of every supplied `waveUrl` with an environment-specific CloudFront base and downloads it using Node's `https` module. | | Amazon EFS | `/mnt/efs/potion-voice//` is the durable model/data/log location and the coupling point between training and synthesis. | | MongoDB | MongoDB Atlas-style `mongodb+srv://...` URIs are selected per message environment. Models represent cloning jobs, profiles, greetings, recordings, and downstream jobs. | | Bugsnag | Both worker entry points initialize Bugsnag with package version, app environment, backend key, and Node release stage. | | Coqui TTS/Trainer | VITS training and inference implementation. The install guide requires a separate editable checkout of Coqui TTS v0.10.2 under ignored `voice-cloning/TTS/`. | | Potion transcription API | `voice-cloning/utils/transcription_utils.py` posts a WAV with a bearer token, then optionally polls for up to 60 seconds. It is used only by the salutation-scoring CLI. Commented examples point at `/api/transcript` on development and staging Potion hosts. | | Dataset sources | Baseline-training instructions retrieve VCTK, LibriTTS, and Potion salutation archives from the private `potion-datasets` S3 bucket. | ## Database & Data Layer Mongoose schemas are defined beside each worker; there is no separate schema package, migration system, repository abstraction, or declared indexes. Most service modules are higher-order factories that bind a Mongoose model and expose basic CRUD methods. Reads commonly add `deleted: false`, while removes are soft deletes. | Model | Role and notable fields | | --- | --- | | `VoiceCloning` | Tracks `userId`, `userAudioProfileId`, `status`, raw `input`, `training_model`, `metadata`, and `deleted`. | | `UserAudioProfile` | Tracks profile `name`, clone `status`, local `training_model_path`, S3 `training_model_s3_path`, and soft deletion. Its schema/service is duplicated in both workers. | | `Salutations` | Caches synthesized audio by `userId`, `userAudioProfileId`, and `firstName`; stores the S3 URL in the historically named `salutationVideo` field. | | `recording_salutations` | Connects a generated greeting to master/dynamic recordings and tracks processing state and derived media URLs. | | `Recordings` | A broad schema shared with the video product. This worker mainly reads original/master video URLs, crop timestamp, user, and dynamic-video type. | | `Job` | Creates the downstream `ai-job` record with recording/user/salutation IDs and a mixed `metadata` payload. | All schemas enable timestamps. Several cross-service payloads and model-asset maps use `Schema.Types.Mixed`, so MongoDB does not enforce their internal shape. ## Connectivity & Configuration The PM2 YAML files are the only environment templates. In this checkout sensitive values are redacted; production values should remain secret rather than being committed. | Variable | Purpose | | --- | --- | | `SQS_URL` | Queue consumed by the current worker. Checked-in examples use environment-specific FIFO queues in `us-west-2`. | | `MONGODB_URI_DEV`, `MONGODB_URI_STAGING`, `MONGODB_URI_PROD` | MongoDB URI selected from the **message's** `env`. Not every PM2 file supplies all three. | | `POTION_APP_ENV` | Used by worker code in the Bugsnag app-version string and by the shared Bugsnag helper. | | `NODE_ENV` | Bugsnag `releaseStage`; PM2 sets it to `production` even in the synthesis development config. | | `BUGSNAG_BACKEND_KEY` | Bugsnag API key. | | `CLOUDFRONT_URL_DEV`, `CLOUDFRONT_URL_STAGING`, `CLOUDFRONT_URL_PROD` | Cloning worker's replacement host for input WAV downloads. | | `APP_ENV` | Present in synthesis PM2 files, but the JavaScript reads `POTION_APP_ENV` instead. | | `TRANSCRIPTION_API_ENDPOINT`, `TRANSCRIPTION_API_TOKEN` | Required only by `score_salutation.py`; token is sent as bearer authentication. | There is no listening application port. TensorBoard is optional and documented on port 6006. Runtime AWS access relies on SDK/CLI credentials or an instance role. Shell tools include `python3`, `tar`, `ffmpeg`, and, for setup, `git`, `unzip`, and `aws`. ## Key Entry Points 1. `voice-cloning-job-handler/index.js` — complete training-worker control flow and its SQS message shape. 2. `voice-synthsizer-job-handler/index.js` — inference worker and handoff to the video job pipeline. 3. `voice-cloning/prepare_datasets.py` — exact input archive layout, sampling conversion, and embedding generation. 4. `voice-cloning/clone_voice.py` — per-speaker VITS fine-tuning configuration. 5. `voice-cloning/synthesize_speech.py` and `voice-cloning/utils/synthesize_utils.py` — inference and 48 kHz WAV production. 6. `voice-cloning/train_multispeaker_baseline_model.py` plus `train_config.py` — shared baseline datasets and model hyperparameters. 7. `voice-cloning/docs/potion-voice-cloning_Installation_Guide.md` — machine sizing, CUDA/system packages, dataset setup, and CLI examples. 8. `app/services/sqs/sqs_service.js` and `app/services/s3/index.js` — shared cloud I/O behavior. ## Notes & Gotchas - A clean clone is not runnable end to end. `voice-cloning/TTS/`, `voice-cloning/pretrained-models/`, generated results, and deployment `app-scripts/` referenced by npm scripts are absent/ignored. The training worker specifically assumes `checkpoint_365000.pth`, then assumes cloning creates `checkpoint_365200.pth` in a directory whose name contains `vits_potion_clone`. - Queue delivery is effectively **at most once**: both workers delete an SQS message before Mongo access, Python execution, or S3 upload. A crash or processing error cannot be retried from that receipt, and no dead-letter handling appears here. - Inference reads EFS-local paths from Mongo, not the uploaded S3 asset map. Training and synthesis hosts therefore need the same `/mnt/efs/potion-voice` mount and path layout. - Training uploads pass `potion-voice-users-training-model/` as the S3 `Bucket` value. Standard S3 bucket names cannot contain `/`; verify whether the environment was intended as a key prefix before relying on this path. - Several commands are assembled as shell strings from message values (`directoryName`, paths, and especially `text`). Quotes or shell metacharacters can break execution and untrusted input would create command-injection risk. - Child-process paths are relative to the worker's current directory (`../voice-cloning/...`), while some Python assets are also opened by relative path. Starting PM2 from a different working directory can therefore break script, encoder, or checkpoint discovery. - Temporary data is only partially cleaned: training archives/extracted files remain under `/tmp`, and synthesis removes the selected 48 kHz file but leaves the original WAV and UUID directory. - Mongo connection retries recursively call `connectDB` without settling the original promise; after an initial connection failure a worker can remain stuck. The selected full Mongo URI is also printed to logs. - `UserAudioProfile.find()` returns an array, but the synthesis worker tests only whether the array is truthy before dereferencing element zero. An empty result follows the exception path rather than the intended “model not found” branch. - PM2 configuration and code use inconsistent environment names (`APP_ENV` versus `POTION_APP_ENV`); the synthesis development file also targets a staging queue while labeling `APP_ENV` as development. The cloning staging CloudFront value is blank in the checked-in example. - Dataset configuration has drift: `train_config.py` overwrites the `POTION_SALUT_*` constants with voice-cloning values, `prepare_datasets.py` advertises a `DAPS` preset but does not implement its branch, and the guide shows some argument values that no longer match argparse choices. - The root manifest declares `index.js` as its main file, but no root `index.js` exists. Worker deployment scripts reference an absent `app-scripts/` tree, and there is no standard `start` or `test` script. - Shared/duplicated code has stale paths: `app/services/voice_cloning/` duplicates the handler implementation, the shared Bugsnag and delete-file utilities are not used by the worker entry points, and `fetchS3Object()` references an undefined `stringifyObj` logger if called. - The install guide pins Coqui TTS v0.10.2 while the Python requirement variants and CUDA guidance span multiple PyTorch/CUDA combinations. Reproduce the intended image deliberately; do not assume the latest packages are compatible.