18 KiB
Potion Voice — Overview
An asynchronous voice-cloning and text-to-speech service for Potion's personalized-video pipeline, combining Node.js queue workers with a GPU-oriented Coqui VITS training and inference toolkit.
Purpose
Potion Voice has no HTTP server or user interface. It provides two continuously running workers: one fine-tunes a per-user voice model from uploaded recordings, and one uses that model to synthesize a personalized greeting and enqueue downstream video-compositing work. The repository also contains Python command-line tools for preparing speech datasets, training the shared multi-speaker baseline, cloning and minimizing individual voices, synthesizing speech, and scoring model or salutation quality.
Tech Stack
| Layer | Technology |
|---|---|
| Worker runtime | Node.js, CommonJS modules; no Node version is declared |
| Process management | PM2, one process per worker |
| ML runtime | Python 3 (the guide targets 3.10), PyTorch, Coqui TTS/Trainer |
| Speech model | VITS with 512-dimensional speaker d-vectors; 22,050 Hz training/inference output |
| Audio processing | Coqui resampling/embedding tools, ffmpeg for 48 kHz output, espeak-ng as the documented phoneme backend |
| Database | MongoDB through Mongoose 6.x |
| Queue and object storage | AWS SDK v2, SQS, S3, CloudFront-hosted source audio |
| Compute and filesystem | GPU-backed EC2 is the documented target; trained assets and logs are placed on an EFS mount |
| Monitoring | Bugsnag for worker exceptions; TensorBoard/TensorBoardX for training runs |
| Evaluation | Resemblyzer speaker similarity, textdistance, and Potion's internal transcription API |
| Tests | No automated test framework, test files, lint command, or CI configuration is present |
Python dependency sets are split across requirements*.txt: development pins PyTorch 1.12.1/CUDA 11.6, the legacy/default set pins PyTorch 1.9.1/CUDA 11.1, production has separate CPU and unpinned-GPU variants, and local development leaves PyTorch unpinned. Every set also installs a private potion-voice-utils Git dependency, although this checkout has no direct import from it.
Directory Structure
.
├── app/services/ Shared Node.js helpers
│ ├── s3/ S3 upload/download wrapper
│ ├── sqs/ SQS receive/delete/send wrapper
│ ├── utils/ Error serialization, Bugsnag helper, file deletion
│ └── voice_cloning/ Older duplicate VoiceCloning model/service
├── voice-cloning-job-handler/ Per-user model-training worker
│ ├── index.js Queue loop and end-to-end orchestration
│ ├── user_audio_profile/ Mongoose schema and CRUD service
│ ├── voice_cloning/ Mongoose schema and CRUD service
│ └── pm2-{development,production}.yml
├── voice-synthsizer-job-handler/ Greeting-synthesis worker (directory typo is historical)
│ ├── index.js Queue loop, synthesis, upload, downstream job creation
│ ├── job/ Downstream AI job schema/service
│ ├── recording/ Large shared Recording schema
│ ├── recording_salutation/ Dynamic-video salutation schema
│ ├── salutation/ Reusable generated-salutation schema/service
│ ├── user_audio_profile/ Duplicate profile schema/service
│ └── pm2-{development,production}.yml
├── voice-cloning/ Python ML and audio toolkit
│ ├── assets/ Speaker encoder and World Gender Name Dictionary data
│ ├── docs/ EC2 setup and command examples
│ ├── utils/ Synthesis, similarity, name matching, transcription helpers
│ ├── prepare_datasets.py Archive extraction, resampling, d-vector generation
│ ├── train_multispeaker_baseline_model.py
│ ├── clone_voice.py Fine-tunes the baseline for one speaker
│ ├── minimize_cloned_voice_model.py Removes training-only model state
│ ├── synthesize_speech.py Generates and resamples a WAV
│ └── score_*.py Manual model/salutation evaluation tools
├── requirements*.txt Python environment variants
├── package.json Shared/root Node dependencies
└── README.md One-line project description
This is not configured as an npm workspace. There are three package manifests with largely duplicated dependencies; the worker code resolves shared modules and, depending on installation layout, dependencies from the repository root.
Architecture
Queue contracts
| Worker | Expected SQS message body |
|---|---|
| Voice cloning | JSON with job._doc._id, job._doc.userAudioProfileId, job._doc.metadata.directoryName, job._doc.input[], and top-level job.env. Each input item contains waveUrl and originalText. |
| Synthesis | JSON with userAudioProfileId, text, firstName, salutationId, recordingId, baseUrlForPotionAi, and env. |
In both workers, the message's env selects the Mongo URI and environment-specific storage resources. This is separate from the process-level environment used to configure PM2 and Bugsnag.
Voice-cloning flow
voice-cloning-job-handler/index.jsshort-polls one message from the configured SQS FIFO queue and immediately deletes it.- It selects a MongoDB connection and CloudFront base URL from the message environment, then marks both the
VoiceCloningandUserAudioProfiledocuments asprocessing. - It rewrites each recording URL's host to the selected CloudFront host, downloads WAV files over HTTPS, and writes a VCTK-style dataset under
/tmp/<directoryName>/{wav48,txt}/1/. Files are numbered1_001,1_002, and so on. - It archives the dataset and invokes three Python programs as child processes:
prepare_datasets.pycomputes speaker embeddings at 16 kHz, then restores and resamples the training audio to 22,050 Hz.clone_voice.pyfine-tunes the hard-codedpretrained-models/checkpoint_365000.pthVITS baseline. Defaults are batch size 96, 200 epochs, mixed precision, two evaluation samples, and checkpoints every 200 steps.minimize_cloned_voice_model.pyreloadscheckpoint_365200.pth, drops the discriminator and optimizer state, and creates_light.pthplusconfig_light.jsoninference assets.
- Generated datasets, checkpoints, configs, embeddings, and command logs live under
/mnt/efs/potion-voice/<env>/<directoryName>/. Mongo status moves tocompleted, andUserAudioProfile.training_model_pathrecords five local paths (full/light model, full/light config, and speaker embeddings). - The same five files are uploaded through S3 and their returned locations are stored in
training_model_s3_path. The code constructs the bucket argument aspotion-voice-users-training-model/<env>and object keys as<directoryName>/<basename>.
An exception after Mongo connects marks both records error and reports to Bugsnag. There is no compensating queue retry because receipt deletion happens before processing.
Greeting-synthesis flow
voice-synthsizer-job-handler/index.jsreceives and immediately deletes one SQS message, connects to the Mongo database selected byjob.env, and finds a completedUserAudioProfile.- It reads the local EFS paths from
training_model_path;training_model_s3_pathis not used for inference.synthesize_speech.pyloads the light VITS model and the profile's single-speaker embeddings, writes a native-rate WAV, and runsffmpegto create the default 48,000 Hz WAV. - The resampled file is uploaded to bucket
recordings-<env>with a generated key ending in_salutation_<firstName>.wav. - The worker upserts a reusable
Salutationsrecord keyed by user, audio profile, and first name; updates the requestedrecording_salutationsrecord; and loads the associatedRecordingsdocument. - It inserts a new
Job(default typeai-job) containing the original video/greeting, crop timestamp, synthesized greeting URL, request origin, environment, recording IDs, and dynamic-video type. Another service is expected to consume this Mongo-backed job and composite the final personalized video.
Both workers run serially in an infinite loop. Empty polls sleep for two seconds; active queues are processed without that delay. They open and close Mongoose around each message rather than maintaining a process-wide connection.
Python toolkit
The Python scripts are also usable independently from voice-cloning/:
- Baseline training combines VCTK 0.92, LibriTTS train-clean-360, and Potion salutation recordings into a multi-speaker VITS model. The checked-in configuration targets 22,050 Hz audio and 512-dimensional d-vectors. The guide estimates 5–7 days for 100 epochs on an AWS
g5.2xlarge. - Per-user cloning expects matching transcripts and recordings in
txt/1/andwav48/1/; the guide recommends 30 samples and says a default clone takes about one hour ong5.2xlarge. score_cloned_voice.pyandscore_models.pysynthesize fixed sentences and compare Resemblyzer embeddings against real recordings; the latter ranks checkpoint files and reports a top five.score_salutation.pytranscribes a WAV, extracts candidate names, validates them against the included World Gender Name Dictionary, and combines transcription confidence with Jaro-Winkler, Levenshtein, and Match Rating Approach similarity.
Integrations
| Integration | Use and code location |
|---|---|
AWS SQS (us-west-2) |
Environment-specific FIFO queues feed both workers. Shared wrappers are in app/services/sqs/; queue URLs are supplied by PM2 configuration. |
| AWS S3 | app/services/s3/index.js uploads trained model assets and synthesized greetings. AWS credentials are not explicit variables; the AWS SDK's normal credential chain is assumed. |
| CloudFront/HTTPS | The cloning worker replaces the host of every supplied waveUrl with an environment-specific CloudFront base and downloads it using Node's https module. |
| Amazon EFS | /mnt/efs/potion-voice/<env>/<directoryName> is the durable model/data/log location and the coupling point between training and synthesis. |
| MongoDB | MongoDB Atlas-style mongodb+srv://... URIs are selected per message environment. Models represent cloning jobs, profiles, greetings, recordings, and downstream jobs. |
| Bugsnag | Both worker entry points initialize Bugsnag with package version, app environment, backend key, and Node release stage. |
| Coqui TTS/Trainer | VITS training and inference implementation. The install guide requires a separate editable checkout of Coqui TTS v0.10.2 under ignored voice-cloning/TTS/. |
| Potion transcription API | voice-cloning/utils/transcription_utils.py posts a WAV with a bearer token, then optionally polls for up to 60 seconds. It is used only by the salutation-scoring CLI. Commented examples point at /api/transcript on development and staging Potion hosts. |
| Dataset sources | Baseline-training instructions retrieve VCTK, LibriTTS, and Potion salutation archives from the private potion-datasets S3 bucket. |
Database & Data Layer
Mongoose schemas are defined beside each worker; there is no separate schema package, migration system, repository abstraction, or declared indexes. Most service modules are higher-order factories that bind a Mongoose model and expose basic CRUD methods. Reads commonly add deleted: false, while removes are soft deletes.
| Model | Role and notable fields |
|---|---|
VoiceCloning |
Tracks userId, userAudioProfileId, status, raw input, training_model, metadata, and deleted. |
UserAudioProfile |
Tracks profile name, clone status, local training_model_path, S3 training_model_s3_path, and soft deletion. Its schema/service is duplicated in both workers. |
Salutations |
Caches synthesized audio by userId, userAudioProfileId, and firstName; stores the S3 URL in the historically named salutationVideo field. |
recording_salutations |
Connects a generated greeting to master/dynamic recordings and tracks processing state and derived media URLs. |
Recordings |
A broad schema shared with the video product. This worker mainly reads original/master video URLs, crop timestamp, user, and dynamic-video type. |
Job |
Creates the downstream ai-job record with recording/user/salutation IDs and a mixed metadata payload. |
All schemas enable timestamps. Several cross-service payloads and model-asset maps use Schema.Types.Mixed, so MongoDB does not enforce their internal shape.
Connectivity & Configuration
The PM2 YAML files are the only environment templates. In this checkout sensitive values are redacted; production values should remain secret rather than being committed.
| Variable | Purpose |
|---|---|
SQS_URL |
Queue consumed by the current worker. Checked-in examples use environment-specific FIFO queues in us-west-2. |
MONGODB_URI_DEV, MONGODB_URI_STAGING, MONGODB_URI_PROD |
MongoDB URI selected from the message's env. Not every PM2 file supplies all three. |
POTION_APP_ENV |
Used by worker code in the Bugsnag app-version string and by the shared Bugsnag helper. |
NODE_ENV |
Bugsnag releaseStage; PM2 sets it to production even in the synthesis development config. |
BUGSNAG_BACKEND_KEY |
Bugsnag API key. |
CLOUDFRONT_URL_DEV, CLOUDFRONT_URL_STAGING, CLOUDFRONT_URL_PROD |
Cloning worker's replacement host for input WAV downloads. |
APP_ENV |
Present in synthesis PM2 files, but the JavaScript reads POTION_APP_ENV instead. |
TRANSCRIPTION_API_ENDPOINT, TRANSCRIPTION_API_TOKEN |
Required only by score_salutation.py; token is sent as bearer authentication. |
There is no listening application port. TensorBoard is optional and documented on port 6006. Runtime AWS access relies on SDK/CLI credentials or an instance role. Shell tools include python3, tar, ffmpeg, and, for setup, git, unzip, and aws.
Key Entry Points
voice-cloning-job-handler/index.js— complete training-worker control flow and its SQS message shape.voice-synthsizer-job-handler/index.js— inference worker and handoff to the video job pipeline.voice-cloning/prepare_datasets.py— exact input archive layout, sampling conversion, and embedding generation.voice-cloning/clone_voice.py— per-speaker VITS fine-tuning configuration.voice-cloning/synthesize_speech.pyandvoice-cloning/utils/synthesize_utils.py— inference and 48 kHz WAV production.voice-cloning/train_multispeaker_baseline_model.pyplustrain_config.py— shared baseline datasets and model hyperparameters.voice-cloning/docs/potion-voice-cloning_Installation_Guide.md— machine sizing, CUDA/system packages, dataset setup, and CLI examples.app/services/sqs/sqs_service.jsandapp/services/s3/index.js— shared cloud I/O behavior.
Notes & Gotchas
- A clean clone is not runnable end to end.
voice-cloning/TTS/,voice-cloning/pretrained-models/, generated results, and deploymentapp-scripts/referenced by npm scripts are absent/ignored. The training worker specifically assumescheckpoint_365000.pth, then assumes cloning createscheckpoint_365200.pthin a directory whose name containsvits_potion_clone. - Queue delivery is effectively at most once: both workers delete an SQS message before Mongo access, Python execution, or S3 upload. A crash or processing error cannot be retried from that receipt, and no dead-letter handling appears here.
- Inference reads EFS-local paths from Mongo, not the uploaded S3 asset map. Training and synthesis hosts therefore need the same
/mnt/efs/potion-voicemount and path layout. - Training uploads pass
potion-voice-users-training-model/<env>as the S3Bucketvalue. Standard S3 bucket names cannot contain/; verify whether the environment was intended as a key prefix before relying on this path. - Several commands are assembled as shell strings from message values (
directoryName, paths, and especiallytext). Quotes or shell metacharacters can break execution and untrusted input would create command-injection risk. - Child-process paths are relative to the worker's current directory (
../voice-cloning/...), while some Python assets are also opened by relative path. Starting PM2 from a different working directory can therefore break script, encoder, or checkpoint discovery. - Temporary data is only partially cleaned: training archives/extracted files remain under
/tmp, and synthesis removes the selected 48 kHz file but leaves the original WAV and UUID directory. - Mongo connection retries recursively call
connectDBwithout settling the original promise; after an initial connection failure a worker can remain stuck. The selected full Mongo URI is also printed to logs. UserAudioProfile.find()returns an array, but the synthesis worker tests only whether the array is truthy before dereferencing element zero. An empty result follows the exception path rather than the intended “model not found” branch.- PM2 configuration and code use inconsistent environment names (
APP_ENVversusPOTION_APP_ENV); the synthesis development file also targets a staging queue while labelingAPP_ENVas development. The cloning staging CloudFront value is blank in the checked-in example. - Dataset configuration has drift:
train_config.pyoverwrites thePOTION_SALUT_*constants with voice-cloning values,prepare_datasets.pyadvertises aDAPSpreset but does not implement its branch, and the guide shows some argument values that no longer match argparse choices. - The root manifest declares
index.jsas its main file, but no rootindex.jsexists. Worker deployment scripts reference an absentapp-scripts/tree, and there is no standardstartortestscript. - Shared/duplicated code has stale paths:
app/services/voice_cloning/duplicates the handler implementation, the shared Bugsnag and delete-file utilities are not used by the worker entry points, andfetchS3Object()references an undefinedstringifyObjlogger if called. - The install guide pins Coqui TTS v0.10.2 while the Python requirement variants and CUDA guidance span multiple PyTorch/CUDA combinations. Reproduce the intended image deliberately; do not assume the latest packages are compatible.