Files
project-work/sources/OVERVIEW.md
2026-09-11 10:54:37 -04:00

18 KiB
Raw Permalink Blame History

Potion Voice — Overview

An asynchronous voice-cloning and text-to-speech service for Potion's personalized-video pipeline, combining Node.js queue workers with a GPU-oriented Coqui VITS training and inference toolkit.

Purpose

Potion Voice has no HTTP server or user interface. It provides two continuously running workers: one fine-tunes a per-user voice model from uploaded recordings, and one uses that model to synthesize a personalized greeting and enqueue downstream video-compositing work. The repository also contains Python command-line tools for preparing speech datasets, training the shared multi-speaker baseline, cloning and minimizing individual voices, synthesizing speech, and scoring model or salutation quality.

Tech Stack

Layer Technology
Worker runtime Node.js, CommonJS modules; no Node version is declared
Process management PM2, one process per worker
ML runtime Python 3 (the guide targets 3.10), PyTorch, Coqui TTS/Trainer
Speech model VITS with 512-dimensional speaker d-vectors; 22,050 Hz training/inference output
Audio processing Coqui resampling/embedding tools, ffmpeg for 48 kHz output, espeak-ng as the documented phoneme backend
Database MongoDB through Mongoose 6.x
Queue and object storage AWS SDK v2, SQS, S3, CloudFront-hosted source audio
Compute and filesystem GPU-backed EC2 is the documented target; trained assets and logs are placed on an EFS mount
Monitoring Bugsnag for worker exceptions; TensorBoard/TensorBoardX for training runs
Evaluation Resemblyzer speaker similarity, textdistance, and Potion's internal transcription API
Tests No automated test framework, test files, lint command, or CI configuration is present

Python dependency sets are split across requirements*.txt: development pins PyTorch 1.12.1/CUDA 11.6, the legacy/default set pins PyTorch 1.9.1/CUDA 11.1, production has separate CPU and unpinned-GPU variants, and local development leaves PyTorch unpinned. Every set also installs a private potion-voice-utils Git dependency, although this checkout has no direct import from it.

Directory Structure

.
├── app/services/                       Shared Node.js helpers
│   ├── s3/                             S3 upload/download wrapper
│   ├── sqs/                            SQS receive/delete/send wrapper
│   ├── utils/                          Error serialization, Bugsnag helper, file deletion
│   └── voice_cloning/                  Older duplicate VoiceCloning model/service
├── voice-cloning-job-handler/          Per-user model-training worker
│   ├── index.js                        Queue loop and end-to-end orchestration
│   ├── user_audio_profile/             Mongoose schema and CRUD service
│   ├── voice_cloning/                  Mongoose schema and CRUD service
│   └── pm2-{development,production}.yml
├── voice-synthsizer-job-handler/       Greeting-synthesis worker (directory typo is historical)
│   ├── index.js                        Queue loop, synthesis, upload, downstream job creation
│   ├── job/                            Downstream AI job schema/service
│   ├── recording/                      Large shared Recording schema
│   ├── recording_salutation/           Dynamic-video salutation schema
│   ├── salutation/                     Reusable generated-salutation schema/service
│   ├── user_audio_profile/             Duplicate profile schema/service
│   └── pm2-{development,production}.yml
├── voice-cloning/                      Python ML and audio toolkit
│   ├── assets/                         Speaker encoder and World Gender Name Dictionary data
│   ├── docs/                           EC2 setup and command examples
│   ├── utils/                          Synthesis, similarity, name matching, transcription helpers
│   ├── prepare_datasets.py             Archive extraction, resampling, d-vector generation
│   ├── train_multispeaker_baseline_model.py
│   ├── clone_voice.py                  Fine-tunes the baseline for one speaker
│   ├── minimize_cloned_voice_model.py  Removes training-only model state
│   ├── synthesize_speech.py            Generates and resamples a WAV
│   └── score_*.py                      Manual model/salutation evaluation tools
├── requirements*.txt                   Python environment variants
├── package.json                        Shared/root Node dependencies
└── README.md                           One-line project description

This is not configured as an npm workspace. There are three package manifests with largely duplicated dependencies; the worker code resolves shared modules and, depending on installation layout, dependencies from the repository root.

Architecture

Queue contracts

Worker Expected SQS message body
Voice cloning JSON with job._doc._id, job._doc.userAudioProfileId, job._doc.metadata.directoryName, job._doc.input[], and top-level job.env. Each input item contains waveUrl and originalText.
Synthesis JSON with userAudioProfileId, text, firstName, salutationId, recordingId, baseUrlForPotionAi, and env.

In both workers, the message's env selects the Mongo URI and environment-specific storage resources. This is separate from the process-level environment used to configure PM2 and Bugsnag.

Voice-cloning flow

  1. voice-cloning-job-handler/index.js short-polls one message from the configured SQS FIFO queue and immediately deletes it.
  2. It selects a MongoDB connection and CloudFront base URL from the message environment, then marks both the VoiceCloning and UserAudioProfile documents as processing.
  3. It rewrites each recording URL's host to the selected CloudFront host, downloads WAV files over HTTPS, and writes a VCTK-style dataset under /tmp/<directoryName>/{wav48,txt}/1/. Files are numbered 1_001, 1_002, and so on.
  4. It archives the dataset and invokes three Python programs as child processes:
    • prepare_datasets.py computes speaker embeddings at 16 kHz, then restores and resamples the training audio to 22,050 Hz.
    • clone_voice.py fine-tunes the hard-coded pretrained-models/checkpoint_365000.pth VITS baseline. Defaults are batch size 96, 200 epochs, mixed precision, two evaluation samples, and checkpoints every 200 steps.
    • minimize_cloned_voice_model.py reloads checkpoint_365200.pth, drops the discriminator and optimizer state, and creates _light.pth plus config_light.json inference assets.
  5. Generated datasets, checkpoints, configs, embeddings, and command logs live under /mnt/efs/potion-voice/<env>/<directoryName>/. Mongo status moves to completed, and UserAudioProfile.training_model_path records five local paths (full/light model, full/light config, and speaker embeddings).
  6. The same five files are uploaded through S3 and their returned locations are stored in training_model_s3_path. The code constructs the bucket argument as potion-voice-users-training-model/<env> and object keys as <directoryName>/<basename>.

An exception after Mongo connects marks both records error and reports to Bugsnag. There is no compensating queue retry because receipt deletion happens before processing.

Greeting-synthesis flow

  1. voice-synthsizer-job-handler/index.js receives and immediately deletes one SQS message, connects to the Mongo database selected by job.env, and finds a completed UserAudioProfile.
  2. It reads the local EFS paths from training_model_path; training_model_s3_path is not used for inference. synthesize_speech.py loads the light VITS model and the profile's single-speaker embeddings, writes a native-rate WAV, and runs ffmpeg to create the default 48,000 Hz WAV.
  3. The resampled file is uploaded to bucket recordings-<env> with a generated key ending in _salutation_<firstName>.wav.
  4. The worker upserts a reusable Salutations record keyed by user, audio profile, and first name; updates the requested recording_salutations record; and loads the associated Recordings document.
  5. It inserts a new Job (default type ai-job) containing the original video/greeting, crop timestamp, synthesized greeting URL, request origin, environment, recording IDs, and dynamic-video type. Another service is expected to consume this Mongo-backed job and composite the final personalized video.

Both workers run serially in an infinite loop. Empty polls sleep for two seconds; active queues are processed without that delay. They open and close Mongoose around each message rather than maintaining a process-wide connection.

Python toolkit

The Python scripts are also usable independently from voice-cloning/:

  • Baseline training combines VCTK 0.92, LibriTTS train-clean-360, and Potion salutation recordings into a multi-speaker VITS model. The checked-in configuration targets 22,050 Hz audio and 512-dimensional d-vectors. The guide estimates 5–7 days for 100 epochs on an AWS g5.2xlarge.
  • Per-user cloning expects matching transcripts and recordings in txt/1/ and wav48/1/; the guide recommends 30 samples and says a default clone takes about one hour on g5.2xlarge.
  • score_cloned_voice.py and score_models.py synthesize fixed sentences and compare Resemblyzer embeddings against real recordings; the latter ranks checkpoint files and reports a top five.
  • score_salutation.py transcribes a WAV, extracts candidate names, validates them against the included World Gender Name Dictionary, and combines transcription confidence with Jaro-Winkler, Levenshtein, and Match Rating Approach similarity.

Integrations

Integration Use and code location
AWS SQS (us-west-2) Environment-specific FIFO queues feed both workers. Shared wrappers are in app/services/sqs/; queue URLs are supplied by PM2 configuration.
AWS S3 app/services/s3/index.js uploads trained model assets and synthesized greetings. AWS credentials are not explicit variables; the AWS SDK's normal credential chain is assumed.
CloudFront/HTTPS The cloning worker replaces the host of every supplied waveUrl with an environment-specific CloudFront base and downloads it using Node's https module.
Amazon EFS /mnt/efs/potion-voice/<env>/<directoryName> is the durable model/data/log location and the coupling point between training and synthesis.
MongoDB MongoDB Atlas-style mongodb+srv://... URIs are selected per message environment. Models represent cloning jobs, profiles, greetings, recordings, and downstream jobs.
Bugsnag Both worker entry points initialize Bugsnag with package version, app environment, backend key, and Node release stage.
Coqui TTS/Trainer VITS training and inference implementation. The install guide requires a separate editable checkout of Coqui TTS v0.10.2 under ignored voice-cloning/TTS/.
Potion transcription API voice-cloning/utils/transcription_utils.py posts a WAV with a bearer token, then optionally polls for up to 60 seconds. It is used only by the salutation-scoring CLI. Commented examples point at /api/transcript on development and staging Potion hosts.
Dataset sources Baseline-training instructions retrieve VCTK, LibriTTS, and Potion salutation archives from the private potion-datasets S3 bucket.

Database & Data Layer

Mongoose schemas are defined beside each worker; there is no separate schema package, migration system, repository abstraction, or declared indexes. Most service modules are higher-order factories that bind a Mongoose model and expose basic CRUD methods. Reads commonly add deleted: false, while removes are soft deletes.

Model Role and notable fields
VoiceCloning Tracks userId, userAudioProfileId, status, raw input, training_model, metadata, and deleted.
UserAudioProfile Tracks profile name, clone status, local training_model_path, S3 training_model_s3_path, and soft deletion. Its schema/service is duplicated in both workers.
Salutations Caches synthesized audio by userId, userAudioProfileId, and firstName; stores the S3 URL in the historically named salutationVideo field.
recording_salutations Connects a generated greeting to master/dynamic recordings and tracks processing state and derived media URLs.
Recordings A broad schema shared with the video product. This worker mainly reads original/master video URLs, crop timestamp, user, and dynamic-video type.
Job Creates the downstream ai-job record with recording/user/salutation IDs and a mixed metadata payload.

All schemas enable timestamps. Several cross-service payloads and model-asset maps use Schema.Types.Mixed, so MongoDB does not enforce their internal shape.

Connectivity & Configuration

The PM2 YAML files are the only environment templates. In this checkout sensitive values are redacted; production values should remain secret rather than being committed.

Variable Purpose
SQS_URL Queue consumed by the current worker. Checked-in examples use environment-specific FIFO queues in us-west-2.
MONGODB_URI_DEV, MONGODB_URI_STAGING, MONGODB_URI_PROD MongoDB URI selected from the message's env. Not every PM2 file supplies all three.
POTION_APP_ENV Used by worker code in the Bugsnag app-version string and by the shared Bugsnag helper.
NODE_ENV Bugsnag releaseStage; PM2 sets it to production even in the synthesis development config.
BUGSNAG_BACKEND_KEY Bugsnag API key.
CLOUDFRONT_URL_DEV, CLOUDFRONT_URL_STAGING, CLOUDFRONT_URL_PROD Cloning worker's replacement host for input WAV downloads.
APP_ENV Present in synthesis PM2 files, but the JavaScript reads POTION_APP_ENV instead.
TRANSCRIPTION_API_ENDPOINT, TRANSCRIPTION_API_TOKEN Required only by score_salutation.py; token is sent as bearer authentication.

There is no listening application port. TensorBoard is optional and documented on port 6006. Runtime AWS access relies on SDK/CLI credentials or an instance role. Shell tools include python3, tar, ffmpeg, and, for setup, git, unzip, and aws.

Key Entry Points

  1. voice-cloning-job-handler/index.js — complete training-worker control flow and its SQS message shape.
  2. voice-synthsizer-job-handler/index.js — inference worker and handoff to the video job pipeline.
  3. voice-cloning/prepare_datasets.py — exact input archive layout, sampling conversion, and embedding generation.
  4. voice-cloning/clone_voice.py — per-speaker VITS fine-tuning configuration.
  5. voice-cloning/synthesize_speech.py and voice-cloning/utils/synthesize_utils.py — inference and 48 kHz WAV production.
  6. voice-cloning/train_multispeaker_baseline_model.py plus train_config.py — shared baseline datasets and model hyperparameters.
  7. voice-cloning/docs/potion-voice-cloning_Installation_Guide.md — machine sizing, CUDA/system packages, dataset setup, and CLI examples.
  8. app/services/sqs/sqs_service.js and app/services/s3/index.js — shared cloud I/O behavior.

Notes & Gotchas

  • A clean clone is not runnable end to end. voice-cloning/TTS/, voice-cloning/pretrained-models/, generated results, and deployment app-scripts/ referenced by npm scripts are absent/ignored. The training worker specifically assumes checkpoint_365000.pth, then assumes cloning creates checkpoint_365200.pth in a directory whose name contains vits_potion_clone.
  • Queue delivery is effectively at most once: both workers delete an SQS message before Mongo access, Python execution, or S3 upload. A crash or processing error cannot be retried from that receipt, and no dead-letter handling appears here.
  • Inference reads EFS-local paths from Mongo, not the uploaded S3 asset map. Training and synthesis hosts therefore need the same /mnt/efs/potion-voice mount and path layout.
  • Training uploads pass potion-voice-users-training-model/<env> as the S3 Bucket value. Standard S3 bucket names cannot contain /; verify whether the environment was intended as a key prefix before relying on this path.
  • Several commands are assembled as shell strings from message values (directoryName, paths, and especially text). Quotes or shell metacharacters can break execution and untrusted input would create command-injection risk.
  • Child-process paths are relative to the worker's current directory (../voice-cloning/...), while some Python assets are also opened by relative path. Starting PM2 from a different working directory can therefore break script, encoder, or checkpoint discovery.
  • Temporary data is only partially cleaned: training archives/extracted files remain under /tmp, and synthesis removes the selected 48 kHz file but leaves the original WAV and UUID directory.
  • Mongo connection retries recursively call connectDB without settling the original promise; after an initial connection failure a worker can remain stuck. The selected full Mongo URI is also printed to logs.
  • UserAudioProfile.find() returns an array, but the synthesis worker tests only whether the array is truthy before dereferencing element zero. An empty result follows the exception path rather than the intended “model not found” branch.
  • PM2 configuration and code use inconsistent environment names (APP_ENV versus POTION_APP_ENV); the synthesis development file also targets a staging queue while labeling APP_ENV as development. The cloning staging CloudFront value is blank in the checked-in example.
  • Dataset configuration has drift: train_config.py overwrites the POTION_SALUT_* constants with voice-cloning values, prepare_datasets.py advertises a DAPS preset but does not implement its branch, and the guide shows some argument values that no longer match argparse choices.
  • The root manifest declares index.js as its main file, but no root index.js exists. Worker deployment scripts reference an absent app-scripts/ tree, and there is no standard start or test script.
  • Shared/duplicated code has stale paths: app/services/voice_cloning/ duplicates the handler implementation, the shared Bugsnag and delete-file utilities are not used by the worker entry points, and fetchS3Object() references an undefined stringifyObj logger if called.
  • The install guide pins Coqui TTS v0.10.2 while the Python requirement variants and CUDA guidance span multiple PyTorch/CUDA combinations. Reproduce the intended image deliberately; do not assume the latest packages are compatible.