890 lines
37 KiB
Markdown
890 lines
37 KiB
Markdown
1. There are not common confi files. There are files like package.json, train_config.py and pm2-development.yml.
|
||
|
||
2. 'tier' or 'tiers' is not mentioned anywhere in the current code.
|
||
|
||
3. pro-v2 is mentioned in git history. he history shows that pro_v2 was introduced as a voice-cloning tier and later removed from the current checkout. I’m tracing its pipeline behavior now so I can distinguish the label
|
||
from the actual model/configuration it selected.
|
||
|
||
|
||
• pro_v2 was a voice-cloning tier/version identifier found only in a checkpoint branch, not the current files.
|
||
|
||
When set on a job, it:
|
||
|
||
- Marked model records as pro_v2
|
||
- Kept its temporary, EFS, and S3 assets separate from legacy assets
|
||
- Prevented legacy models from satisfying pro_v2 jobs
|
||
|
||
Missing/null tier meant legacy behavior. Notably, it did not select a different training algorithm in the implementation—it mainly provided version tagging and storage isolation.
|
||
|
||
|
||
4. The code and history tell different stories. The implementation spelling was pro_v2.
|
||
|
||
## Current code
|
||
|
||
- No tracked application file contains pro_v2 or pro-v2.
|
||
- The only working-tree mentions are in the untracked holistic-rubric.md:5.
|
||
- The current worker still unconditionally reads job._doc at voice-cloning-job-handler/index.js:104.
|
||
- There is no producer or formal pro_v2 message schema here.
|
||
|
||
## Git history
|
||
|
||
All implementation matches occur in generated raccoon-checkpoint history—not the original named product commits.
|
||
|
||
Commit 97eb393 introduced an extensive tier implementation:
|
||
|
||
- Accepted pro_v2 from job.tier, job._doc.tier, or job._doc.metadata.tier.
|
||
- Added tier and training_model_tier fields to Mongo schemas.
|
||
- Prevented pro_v2 jobs from reusing legacy model assets.
|
||
- Namespaced temporary, EFS, and S3 paths under pro_v2/.
|
||
- Added validation, documentation, and tests.
|
||
- Kept the same Python training algorithm and baseline model.
|
||
|
||
Crucially, that implementation still required the cloning payload itself to be under _doc; its validator explicitly rejected jobs without _doc. Therefore, it did not support the flat
|
||
payload described by the rubric.
|
||
|
||
Commit cf8da75 subsequently removed that tier implementation along with the larger queue-worker refactor and tests, returning the application files to their current state.
|
||
|
||
Commit 512f28c later captured holistic-rubric.md; its pro_v2 mentions are documentation only.
|
||
|
||
5. › what is 'pro_v2' - how is it classified?
|
||
|
||
|
||
• pro_v2 is classified as a voice-cloning job tier identifier—a string label, not a model or algorithm.
|
||
|
||
- Rubric: identifies requests using the newer flat JSON transport shape.
|
||
- Historical code: treated it as a model/artifact compatibility tier, persisted in tier and training_model_tier and used to isolate storage.
|
||
- Current code: does not define or recognize it.
|
||
|
||
There is no evidence that it means a pricing or subscription tier.
|
||
|
||
6. › what is 'pro_v2' - how is it classified?
|
||
|
||
|
||
• pro_v2 is classified as a voice-cloning job tier identifier—a string label, not a model or algorithm.
|
||
|
||
- Rubric: identifies requests using the newer flat JSON transport shape.
|
||
- Historical code: treated it as a model/artifact compatibility tier, persisted in tier and training_model_tier and used to isolate storage.
|
||
- Current code: does not define or recognize it.
|
||
|
||
There is no evidence that it means a pricing or subscription tier.
|
||
|
||
|
||
7. › is tier spec'd somewhere in the code and if not, from the repo history what does it mean?
|
||
|
||
|
||
• No—current tracked code does not specify a voice-cloning tier. The schemas and worker neither define nor read one. Only the untracked rubric mentions it.
|
||
|
||
In checkpoint history (97eb393), tier meant an artifact-compatibility namespace:
|
||
|
||
- null/missing → legacy cloning assets
|
||
- pro_v2 → versioned cloning assets stored under separate pro_v2/ temporary, EFS, and S3 paths
|
||
- Retries could reuse only assets with the matching tier
|
||
- Records stored it as tier and training_model_tier
|
||
|
||
It did not select a different training algorithm. History provides no authoritative business definition for “pro,” and that checkpoint implementation was later removed.
|
||
|
||
8. what other cloning job types or classes are defined in the current codebase?
|
||
• The current codebase defines no cloning-job types, tiers, or class hierarchy.
|
||
|
||
It has:
|
||
|
||
- One VoiceCloning Mongoose model, duplicated in two locations but representing the same record—not separate classes:
|
||
- handler model (voice-cloning-job-handler/voice_cloning/voice_cloning_model.js:4)
|
||
- app model (app/services/voice_cloning/voice_cloning_model.js:4)
|
||
|
||
- One associated UserAudioProfile model.
|
||
- Lifecycle states: created, processing, completed, and error. These are statuses, not job classes.
|
||
- Environment routing: development, staging, and production—not cloning types.
|
||
|
||
There is a separate speech-synthesis worker with a generic Job.type defaulting to ai-job, but that is downstream synthesis, not another cloning job type.
|
||
|
||
Dataset presets such as VCTK, DAPS, and potion_voice_cloning exist in the Python tooling, but the worker hardcodes potion_voice_cloning; they classify datasets, not jobs.
|
||
|
||
9.
|
||
› what's your assessment as to what this codebase does? then, how is it used - as in - objects are created and methods called, some other way. Where in the voice processing workflow
|
||
would this code fit in?
|
||
|
||
• The architecture is clear: it is process-oriented rather than an object-oriented library. Two long-running Node workers coordinate MongoDB/AWS/filesystem state and invoke the ML
|
||
implementation as separate Python command-line programs. I’m confirming the training mode and handoff artifacts before summarizing the workflow.
|
||
|
||
• ## Assessment
|
||
|
||
This is a batch-oriented personalized text-to-speech system, apparently built for inserting spoken names or greetings into personalized videos.
|
||
|
||
It has two production workers:
|
||
|
||
1. A voice-cloning worker fine-tunes a Coqui VITS baseline model using a user’s recordings.
|
||
2. A synthesis worker uses that personalized model to generate a WAV for supplied text, uploads it, and creates a downstream video-processing job.
|
||
|
||
It is not an API server or an object-oriented job framework.
|
||
|
||
## How it runs
|
||
|
||
PM2 starts each index.js as a persistent process. Each file immediately calls init() and enters an infinite SQS polling loop.
|
||
|
||
The main interaction pattern is:
|
||
|
||
- Receive and parse an SQS JSON message.
|
||
- Connect to MongoDB.
|
||
- Call service methods such as update(), find(), and insertMany().
|
||
- Read and write files in /tmp and EFS.
|
||
- Launch Python command-line programs through child_process.exec.
|
||
- Upload generated assets to S3.
|
||
|
||
Mongoose models are bound to service objects at module load time. For example, voice-cloning-job-handler/voice_cloning/index.js:1 effectively creates:
|
||
|
||
VoiceCloningService(VoiceCloningModel)
|
||
|
||
The cloning worker does not create the VoiceCloning or UserAudioProfile records. It assumes an upstream service already created them and supplied their IDs in the queue message. It then
|
||
updates those records through processing, completed, or error.
|
||
|
||
The Python code does instantiate ML objects—Vits, Trainer, and SpeakerManager—but Node invokes those scripts as separate operating-system processes rather than importing them.
|
||
|
||
## Workflow position
|
||
|
||
External application (not in repository)
|
||
├─ collects voice recordings
|
||
├─ creates VoiceCloning + UserAudioProfile records
|
||
└─ sends cloning SQS message
|
||
│
|
||
▼
|
||
Voice-cloning worker
|
||
├─ downloads recordings
|
||
├─ prepares/resamples data and computes speaker embeddings
|
||
├─ fine-tunes the baseline VITS model
|
||
├─ removes training-only model components
|
||
└─ saves model paths in MongoDB and uploads assets to S3
|
||
│
|
||
▼
|
||
Voice-synthesis worker
|
||
├─ receives text such as a person’s name/greeting
|
||
├─ loads the completed personalized model
|
||
├─ generates and resamples a WAV
|
||
├─ uploads the WAV to S3
|
||
└─ creates a generic AI/video-processing job
|
||
│
|
||
▼
|
||
Video composition system (not in repository)
|
||
|
||
The relevant implementation stages are:
|
||
|
||
- Dataset preparation: voice-cloning/prepare_datasets.py:103
|
||
- Per-speaker VITS fine-tuning: voice-cloning/clone_voice.py:156
|
||
- Lightweight inference model creation: voice-cloning/minimize_cloned_voice_model.py:82
|
||
- Speech generation: voice-cloning/synthesize_speech.py:96
|
||
|
||
The separate baseline-training and scoring scripts appear to be offline development/quality tools, not part of the queue-driven production path.
|
||
|
||
In this workflow, the rubric’s pro_v2 problem occurs solely at the cloning worker’s SQS ingress—before database updates or voice processing. It concerns how the incoming object is
|
||
unpacked, not a different cloning algorithm.
|
||
|
||
10. › are there any other voice-cloning job labels? pro_v2 would be one if implemented (and it's not).
|
||
|
||
|
||
• No. There are no implemented voice-cloning job labels in the current code.
|
||
|
||
In checkpoint history:
|
||
|
||
- pro_v2 was the only accepted explicit label.
|
||
- Missing or null tier meant legacy behavior; "legacy" was not an actual label.
|
||
- pro_v3 appeared only in tests as an unsupported value.
|
||
|
||
Strings such as created, processing, completed, and error are statuses, while potion_voice_cloning is a dataset preset—not job labels.
|
||
|
||
|
||
11. › are there any other voice-cloning job labels? pro_v2 would be one if implemented (and it's not).
|
||
|
||
|
||
• No. There are no implemented voice-cloning job labels in the current code.
|
||
|
||
In checkpoint history:
|
||
|
||
- pro_v2 was the only accepted explicit label.
|
||
- Missing or null tier meant legacy behavior; "legacy" was not an actual label.
|
||
- pro_v3 appeared only in tests as an unsupported value.
|
||
|
||
Strings such as created, processing, completed, and error are statuses, while potion_voice_cloning is a dataset preset—not job labels.
|
||
|
||
|
||
› what kind of label/tier/thing is currently processed that pro_v2 or pro_v3 would be analogous too?
|
||
|
||
|
||
• There is no explicit analogous label in the current cloning worker. Every accepted job follows one fixed cloning configuration.
|
||
|
||
The practical counterpart is implicit:
|
||
|
||
- Existing/legacy job: cloning fields are inside job._doc.
|
||
- pro_v2 job per the rubric: the same fields are directly on job.
|
||
|
||
So pro_v2 identifies an upstream request cohort whose transport shape changed—not a current in-repo model class.
|
||
|
||
The worker does distinguish env (development, staging, production), but that controls infrastructure selection, not cloning quality or version. All jobs otherwise use the same hardcoded
|
||
dataset preset and baseline model.
|
||
|
||
12. › take me through the voice-cloning-job-handler code
|
||
|
||
|
||
• The handler is a single, self-starting SQS worker that orchestrates per-user VITS model training. Nearly the entire runtime is in voice-cloning-job-handler/index.js:1.
|
||
|
||
## 1. Startup and dependencies
|
||
|
||
The file imports:
|
||
|
||
- AWS SQS and S3 helpers
|
||
- Mongoose
|
||
- Bugsnag
|
||
- Filesystem and HTTPS modules
|
||
- child_process.exec for running Python
|
||
- Handler-local VoiceCloning and UserAudioProfile services
|
||
|
||
Environment variables provide the queue URL, MongoDB URIs, CloudFront origins, and monitoring configuration.
|
||
|
||
PM2 launches index.js as a continuously restarting, single-instance process named training-model.
|
||
|
||
## 2. Expected queue message
|
||
|
||
The worker expects this approximate shape:
|
||
|
||
{
|
||
"_doc": {
|
||
"_id": "voice-cloning-record-id",
|
||
"userAudioProfileId": "profile-id",
|
||
"metadata": {
|
||
"directoryName": "profile-directory"
|
||
},
|
||
"input": [
|
||
{
|
||
"waveUrl": "https://example.com/sample.wav",
|
||
"originalText": "Text spoken in the sample"
|
||
}
|
||
]
|
||
},
|
||
"env": "staging"
|
||
}
|
||
|
||
At queue processing:89 (voice-cloning-job-handler/index.js:89), it:
|
||
|
||
1. Receives one SQS message.
|
||
2. Parses Body as JSON.
|
||
3. Extracts cloning fields from job._doc.
|
||
4. Extracts env from the top level.
|
||
5. Selects the development, staging, or production MongoDB and CloudFront configuration.
|
||
|
||
This is where a flat pro_v2 payload fails: job._doc is absent, so line 104 throws before any processing begins.
|
||
|
||
## 3. Claiming and tracking the job
|
||
|
||
After connecting to MongoDB, the worker immediately deletes the SQS message at line 130.
|
||
|
||
It then updates two pre-existing MongoDB records:
|
||
|
||
- VoiceCloning → processing
|
||
- UserAudioProfile → processing
|
||
|
||
The worker does not create those records. An upstream service—not present here—must create them and enqueue their identifiers.
|
||
|
||
The imported services are factory-bound wrappers around Mongoose models. The worker calls methods such as:
|
||
|
||
voiceCloningService.update({ _id, status: 'processing' })
|
||
userAudioProfileService.update({
|
||
_id: userAudioProfileId,
|
||
status: 'processing'
|
||
})
|
||
|
||
Both services ultimately use findOneAndUpdate({ _id: data._id }, data).
|
||
|
||
## 4. Building the training dataset
|
||
|
||
For each input recording, the worker creates a structure like:
|
||
|
||
/tmp/<directoryName>/
|
||
├── wav48/1/
|
||
│ ├── 1_001.wav
|
||
│ └── 1_002.wav
|
||
└── txt/1/
|
||
├── 1_001.txt
|
||
└── 1_002.txt
|
||
|
||
It downloads each waveUrl, replacing its original host with the environment’s CloudFront origin, and writes the corresponding originalText.
|
||
|
||
It then archives the directory as /tmp/<directoryName>.tgz.
|
||
|
||
## 5. Preparing the audio
|
||
|
||
The first Python command invokes voice-cloning/prepare_datasets.py:66:
|
||
|
||
prepare_datasets.py
|
||
--dataset_preset potion_voice_cloning
|
||
--dataset_archive_path /tmp/<name>.tgz
|
||
--output_path /mnt/efs/potion-voice/<env>/<name>
|
||
|
||
That script:
|
||
|
||
- Extracts the archive
|
||
- Resamples audio to 16 kHz temporarily
|
||
- Computes 512-dimensional speaker embeddings
|
||
- Restores and resamples the training audio to 22.05 kHz
|
||
- Writes speakers.pth
|
||
|
||
## 6. Cloning the voice
|
||
|
||
The worker invokes voice-cloning/clone_voice.py:49 with:
|
||
|
||
- A fixed baseline VITS checkpoint: checkpoint_365000.pth
|
||
- The prepared speaker dataset
|
||
- speakers.pth
|
||
- An EFS output directory
|
||
|
||
The Python script constructs a Coqui VITS model and Trainer, restores the baseline checkpoint, and calls trainer.fit(). This is per-speaker fine-tuning, rather than zero-shot voice
|
||
cloning.
|
||
|
||
Unless overridden, it trains for up to 200 epochs with a batch size of 96.
|
||
|
||
The referenced baseline checkpoint is not tracked in this repository, so deployment must provide it separately.
|
||
|
||
## 7. Minimizing the resulting model
|
||
|
||
The worker searches the results directory for a name containing vits_potion_clone, then invokes voice-cloning/minimize_cloned_voice_model.py:50.
|
||
|
||
That script:
|
||
|
||
- Loads checkpoint_365200.pth
|
||
- Removes the discriminator and training-only state
|
||
- Saves a smaller checkpoint_365200_light.pth
|
||
- Writes config_light.json
|
||
|
||
The lightweight files are intended for speech inference.
|
||
|
||
## 8. Saving the result
|
||
|
||
The worker records five paths on the audio profile:
|
||
|
||
checkpoint_365200.pth
|
||
config.json
|
||
speakers.pth
|
||
checkpoint_365200_light.pth
|
||
config_light.json
|
||
|
||
It marks both records completed, uploads those assets to an environment-specific S3 bucket, and then stores the resulting S3 URLs on UserAudioProfile.
|
||
|
||
The downstream synthesizer worker later reads the lightweight model, config, and speaker embeddings to generate speech.
|
||
|
||
## 9. Failure and polling behavior
|
||
|
||
Failures inside the processing section cause both records to be marked error and reported to Bugsnag.
|
||
|
||
The outer loop closes MongoDB after every poll and immediately polls again while messages are available; it waits two seconds when the queue is empty.
|
||
|
||
Important operational characteristics:
|
||
|
||
- The SQS message is deleted before training, so a crash or training failure cannot be retried from that message.
|
||
- Invalid envelopes fail before status updates.
|
||
- MongoDB retry handling can leave its original promise unresolved after a failed first attempt.
|
||
- Paths and shell commands interpolate message-provided values without validation.
|
||
- Network downloads have no status, timeout, redirect, or error handling.
|
||
- This file starts polling when imported, making isolated testing difficult.
|
||
|
||
|
||
13. › take me through the voice-synthsizer-job-handler code
|
||
|
||
|
||
• The voice-synthsizer-job-handler is the inference-side worker. It does not clone or train voices; it uses a previously trained voice model to generate a personalized spoken greeting,
|
||
then creates a downstream video-processing job.
|
||
|
||
## 1. Startup
|
||
|
||
The runtime is voice-synthsizer-job-handler/index.js:1.
|
||
|
||
It imports:
|
||
|
||
- AWS SQS and S3 helpers
|
||
- Mongoose
|
||
- Bugsnag
|
||
- The UserAudioProfile service
|
||
- Recording, RecordingSalutation, Salutation, and Job models/services
|
||
- child_process.exec for running Python
|
||
- UUID generation for temporary paths and filenames
|
||
|
||
PM2 launches it as a single continuously restarting process named synthsizer-job.
|
||
|
||
## 2. Expected SQS message
|
||
|
||
Unlike the cloning worker, this worker expects a flat object:
|
||
|
||
{
|
||
"userAudioProfileId": "profile-id",
|
||
"text": "Hey, Sarah!",
|
||
"firstName": "Sarah",
|
||
"salutationId": "recording-salutation-id",
|
||
"recordingId": "recording-id",
|
||
"baseUrlForPotionAi": "https://...",
|
||
"env": "production"
|
||
}
|
||
|
||
There is no _doc access and no tier or job-type discriminator.
|
||
|
||
## 3. Receiving the request
|
||
|
||
At processQueue:58 (voice-synthsizer-job-handler/index.js:58), the worker:
|
||
|
||
1. Fetches one SQS message.
|
||
2. Parses the message body.
|
||
3. Immediately deletes the message.
|
||
4. Extracts the fields above.
|
||
5. Selects a MongoDB URI from env.
|
||
6. Connects to MongoDB.
|
||
|
||
As with the cloning handler, deleting the message before doing the work means failures cannot be retried through that SQS delivery.
|
||
|
||
## 4. Loading the cloned voice
|
||
|
||
The worker queries UserAudioProfile for the supplied ID and requires its status to be completed:
|
||
|
||
userAudioProfileService.find({
|
||
_id: userAudioProfileId,
|
||
status: 'completed'
|
||
})
|
||
|
||
From the first matching profile, it reads:
|
||
|
||
- voice_model_light_path
|
||
- voice_model_config_light_path
|
||
- voice_model_speakers_file_path
|
||
- The profile owner’s userId
|
||
|
||
These are local filesystem paths produced by the cloning worker. Although S3 paths are also stored on the profile, this worker does not download or use them. It therefore assumes the
|
||
trained assets remain accessible through shared storage such as EFS.
|
||
|
||
## 5. Generating speech
|
||
|
||
It creates a unique temporary directory and executes voice-cloning/synthesize_speech.py:50:
|
||
|
||
python3 synthesize_speech.py
|
||
--voice_model_path <light checkpoint>
|
||
--voice_model_config_path <light config>
|
||
--speaker_embeddings_path <speakers.pth>
|
||
--txt "<requested text>"
|
||
--output_path <temporary directory>
|
||
|
||
The Python script:
|
||
|
||
1. Loads the lightweight Coqui VITS model.
|
||
2. Loads the speaker embeddings.
|
||
3. Verifies that the embedding file represents one speaker.
|
||
4. Computes the speaker’s mean embedding.
|
||
5. Synthesizes the requested text.
|
||
6. Saves the original WAV.
|
||
7. Uses FFmpeg to produce a 48 kHz WAV.
|
||
|
||
The Node worker finds the output filename containing sr48000.wav.
|
||
|
||
## 6. Uploading the greeting
|
||
|
||
The generated WAV is uploaded to:
|
||
|
||
s3://recordings-<env>/<uuid>_salutation_<firstName>.wav
|
||
|
||
The resulting S3 URL becomes greetingUploadResponse.
|
||
|
||
## 7. Updating application records
|
||
|
||
The worker calls salutationService.updateOrCreate().
|
||
|
||
That service searches by:
|
||
|
||
- firstName
|
||
- userAudioProfileId
|
||
- userId
|
||
|
||
If a matching Salutation exists, it updates its audio URL. Otherwise, it creates a new Mongoose document.
|
||
|
||
It then loads:
|
||
|
||
- The existing RecordingSalutation identified by salutationId
|
||
- The source Recording identified by recordingId
|
||
|
||
If both are suitable, it writes the new WAV URL to RecordingSalutation.salutationVideo.
|
||
|
||
## 8. Creating the downstream video job
|
||
|
||
The worker builds metadata containing:
|
||
|
||
- The original greeting and video
|
||
- Crop timestamp
|
||
- Generated greeting clip
|
||
- Target first name
|
||
- Recording and salutation IDs
|
||
- Request origin
|
||
- Environment
|
||
- Dynamic-video type
|
||
|
||
It inserts a generic Job model (voice-synthsizer-job-handler/job/job_model.js:3) record through jobService.insertMany().
|
||
|
||
Because no explicit type or status is supplied, the schema defaults are:
|
||
|
||
type: ai-job
|
||
status: created
|
||
|
||
This repository does not contain the consumer that processes that job. Presumably another service sees it and combines the generated greeting audio with the original video.
|
||
|
||
## 9. Resulting workflow
|
||
|
||
Completed UserAudioProfile
|
||
│
|
||
▼
|
||
Synthesis SQS message
|
||
│
|
||
▼
|
||
Load lightweight cloned model from shared storage
|
||
│
|
||
▼
|
||
Generate 48 kHz personalized WAV
|
||
│
|
||
▼
|
||
Upload WAV to S3
|
||
│
|
||
├─ update/create Salutation
|
||
├─ update RecordingSalutation
|
||
└─ create ai-job for downstream video processing
|
||
|
||
## Important implementation observations
|
||
|
||
- There is no voice tier or model-selection logic; the profile’s stored paths determine which model is loaded.
|
||
- The message is deleted before database access or synthesis.
|
||
- find() returns an array, but the code tests only whether the array is truthy; an empty array still passes and makes userAudioProfile[0] undefined.
|
||
- User-provided text is interpolated directly into a shell command.
|
||
- Only the resampled WAV is deleted; the original WAV and temporary directory remain.
|
||
- MongoDB retry handling has the same unresolved-promise problem as the cloning worker.
|
||
- The code reads POTION_APP_ENV, while its PM2 files set APP_ENV, so the Bugsnag application version may begin with undefined.
|
||
|
||
|
||
14. • CloudFront is used only by the voice-cloning worker, during training-recording download.
|
||
|
||
## Active use
|
||
|
||
The worker reads three environment-specific origins at voice-cloning-job-handler/index.js:22:
|
||
|
||
- CLOUDFRONT_URL_DEV
|
||
- CLOUDFRONT_URL_STAGING
|
||
- CLOUDFRONT_URL_PROD
|
||
|
||
For every input waveUrl, updateUrl() replaces the original URL origin while preserving its path and query:
|
||
|
||
https://original-host/path/sample.wav
|
||
↓
|
||
https://configured-cloudfront/path/sample.wav
|
||
|
||
The rewritten URL is downloaded at voice-cloning-job-handler/index.js:164 and becomes training input.
|
||
|
||
## Why
|
||
|
||
The associated PR was named “Updated the cloud front access and code,” from branch fix-bucket-access-for-sentences. Its body contains no detailed explanation, but the code indicates
|
||
that CloudFront was introduced to make source recordings accessible through an environment-controlled CDN/domain instead of their original storage host.
|
||
|
||
It is not used for:
|
||
|
||
- Uploading trained models—the worker uploads those directly to S3.
|
||
- Loading trained models—the synthesizer reads shared local/EFS paths.
|
||
- Uploading synthesized audio—the synthesizer uploads directly to S3.
|
||
- Any Python processing.
|
||
|
||
One apparent configuration issue: both checked-in PM2 files leave CLOUDFRONT_URL_STAGING empty. Unless overridden at deployment, a staging job would produce an invalid download URL.
|
||
|
||
|
||
15. › But why is Cloudfront being used as opposed to some other provider
|
||
|
||
|
||
• The repository does not document a vendor-selection rationale. Technically, the code does not require CloudFront specifically—it simply substitutes a configured HTTPS origin into each
|
||
recording URL. Any CDN or proxy preserving the same paths could work.
|
||
|
||
CloudFront was likely chosen because the surrounding infrastructure is already AWS-based:
|
||
|
||
- Recordings and generated assets use S3.
|
||
- Jobs use SQS.
|
||
- Training uses EFS and likely EC2.
|
||
- CloudFront can expose S3 objects through a stable public/custom domain while keeping direct bucket access restricted.
|
||
|
||
The PR branch name, fix-bucket-access-for-sentences, suggests access to source recordings—not CDN performance—was the immediate motivation. The production setting,
|
||
videoassets.sendpotion.com, is presumably a custom domain backed by CloudFront.
|
||
|
||
So CloudFront is an infrastructure choice, not a voice-processing requirement. No evidence shows that another provider was evaluated or excluded.
|
||
|
||
16. tools
|
||
• Using executable imports/calls as “used” — not comments, PR metadata, or documentation — the current dependency map is:
|
||
|
||
voice-cloning-job-handler
|
||
├─ prepare_datasets.py
|
||
│ ├─ train_config.py
|
||
│ ├─ config_se.json
|
||
│ └─ model_se.pth.tar
|
||
├─ clone_voice.py
|
||
└─ minimize_cloned_voice_model.py
|
||
|
||
voice-synthsizer-job-handler
|
||
└─ synthesize_speech.py
|
||
└─ utils/synthesize_utils.py
|
||
|
||
## Not reached by either production worker
|
||
|
||
These are standalone/manual tools:
|
||
|
||
- voice-cloning/train_multispeaker_baseline_model.py:1 — offline baseline-model training; documented in the installation guide.
|
||
- voice-cloning/score_cloned_voice.py:1 — manual cloned-voice similarity scoring; documented.
|
||
- voice-cloning/score_salutation.py:1 — manual transcription/name scoring; documented.
|
||
- voice-cloning/score_models.py:1 — model-ranking experiment; no current code or user documentation references it.
|
||
|
||
Because these are CLI entry points, the repository cannot show whether someone runs them manually. It only shows that the application does not call them.
|
||
|
||
## Utilities used only by those manual tools
|
||
|
||
These are referenced, but never from the production workers:
|
||
|
||
- utils/matching_utils.py → only score_salutation.py
|
||
- utils/transcription_utils.py → only score_salutation.py
|
||
- utils/scoring_utils.py → only score_cloned_voice.py and score_models.py
|
||
|
||
Therefore, those three utility modules and their dependencies are production-dead unless the manual scoring tools are intentionally retained.
|
||
|
||
utils/synthesize_utils.py is production-active because synthesize_speech.py imports it.
|
||
|
||
## Assets with no code references
|
||
|
||
These two files have no source-code references at all:
|
||
|
||
- assets/wgnd_2_0_sources.csv.gz
|
||
- assets/wgnd_2_0_unique_names_only.csv.gz
|
||
|
||
The third names file, wgnd_2_0_unique_names_only_limited_special_chars.csv, is used only by score_salutation.py, so it is also outside the production path.
|
||
|
||
Bottom line: the strongest unused candidate is score_models.py; the two compressed WGND files are completely unreferenced. The scoring/name/transcription group is isolated manual
|
||
tooling, while the four preparation, cloning, minimization, and synthesis scripts form the active application path.
|
||
|
||
|
||
17. UserAudioProfile
|
||
• My best one-line description: UserAudioProfile is a mutable MongoDB manifest for a user’s reusable cloned voice. It is the handoff point between voice training and speech synthesis.
|
||
|
||
Despite its name, it does not contain the user’s source audio. It contains ownership, readiness status, and pointers to the trained model artifacts.
|
||
|
||
## What it is syntactically
|
||
|
||
It is a Mongoose model, not a JavaScript class, TypeScript type, or queue-job type:
|
||
|
||
const UserAudioProfileSchema = mongoose.Schema({...}, {
|
||
timestamps: true
|
||
})
|
||
|
||
module.exports = mongoose.model(
|
||
'UserAudioProfile',
|
||
UserAudioProfileSchema
|
||
)
|
||
|
||
There are two effectively identical copies:
|
||
|
||
- Cloning-worker model (voice-cloning-job-handler/user_audio_profile/user_audio_profile_model.js:4)
|
||
- Synthesizer-worker model (voice-synthsizer-job-handler/user_audio_profile/user_audio_profile_model.js:4)
|
||
|
||
Each worker is a separate process and compiles its own copy of the same MongoDB model. This looks like duplicated local knowledge of a shared database contract, presumably because the
|
||
workers were intended to deploy independently.
|
||
|
||
Mongoose supplies _id automatically and likely stores documents in its default pluralized collection, useraudioprofiles.
|
||
|
||
## Document shape
|
||
|
||
A representative document would look like:
|
||
|
||
{
|
||
_id: ObjectId("..."),
|
||
userId: ObjectId("..."),
|
||
name: "My voice",
|
||
status: "completed",
|
||
|
||
training_model_path: {
|
||
voice_model_path: "/mnt/efs/.../checkpoint_365200.pth",
|
||
voice_model_config_path: "/mnt/efs/.../config.json",
|
||
voice_model_speakers_file_path: "/mnt/efs/.../speakers.pth",
|
||
voice_model_light_path: "/mnt/efs/.../checkpoint_365200_light.pth",
|
||
voice_model_config_light_path: "/mnt/efs/.../config_light.json"
|
||
},
|
||
|
||
training_model_s3_path: {
|
||
// Same keys, with S3 URLs as values
|
||
},
|
||
|
||
deleted: false,
|
||
createdAt: Date,
|
||
updatedAt: Date
|
||
}
|
||
|
||
Field Apparent meaning
|
||
━━━━━━━━━━━━━━━━━━━━━━━━ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||
userId Owner of the voice profile
|
||
──────────────────────── ───────────────────────────────────────────────────────────
|
||
name User-facing name for the profile; unused by these workers
|
||
──────────────────────── ───────────────────────────────────────────────────────────
|
||
status Training/readiness lifecycle
|
||
──────────────────────── ───────────────────────────────────────────────────────────
|
||
training_model_path Shared local/EFS locations of model artifacts
|
||
──────────────────────── ───────────────────────────────────────────────────────────
|
||
training_model_s3_path Uploaded S3 locations of the same artifacts
|
||
──────────────────────── ───────────────────────────────────────────────────────────
|
||
deleted Soft-deletion marker
|
||
──────────────────────── ───────────────────────────────────────────────────────────
|
||
timestamps Creation and modification times
|
||
|
||
The model has no tier, version, model family, language, sampling rate, or immutable training-run identifier.
|
||
|
||
## How code accesses it
|
||
|
||
The directory’s index.js passes the Mongoose model into a service factory:
|
||
|
||
module.exports = UserAudioProfileService(UserAudioProfile)
|
||
|
||
That exports a plain service object with:
|
||
|
||
create
|
||
insertMany
|
||
read
|
||
find
|
||
update
|
||
remove
|
||
removeMany
|
||
|
||
The methods are closure-bound wrappers over Mongoose operations. For example, update() executes:
|
||
|
||
UserAudioProfileModel.findOneAndUpdate(
|
||
{ _id: data._id },
|
||
data,
|
||
{ new: true }
|
||
)
|
||
|
||
Neither worker normally constructs a profile with new UserAudioProfile(). Although the service exposes create(), there are no current callers. Profile creation happens in an upstream
|
||
application absent from this repository.
|
||
|
||
## Role during cloning
|
||
|
||
The queue message supplies userAudioProfileId. The VoiceCloning record also references that profile:
|
||
|
||
VoiceCloning.userAudioProfileId → UserAudioProfile._id
|
||
|
||
The cloning worker uses the profile as the durable destination for the training result:
|
||
|
||
1. Sets its status to processing.
|
||
2. Trains and minimizes a personalized model.
|
||
3. Sets status: completed.
|
||
4. Writes local/EFS model paths.
|
||
5. Uploads the artifacts and adds their S3 paths.
|
||
6. Sets status: error if processing fails.
|
||
|
||
The separate VoiceCloning document represents the training request or operation. UserAudioProfile represents the resulting reusable voice.
|
||
|
||
## Role during synthesis
|
||
|
||
The synthesizer receives userAudioProfileId and queries for:
|
||
|
||
{
|
||
_id: userAudioProfileId,
|
||
status: 'completed',
|
||
deleted: false
|
||
}
|
||
|
||
It then reads three entries from training_model_path:
|
||
|
||
- Lightweight model checkpoint
|
||
- Lightweight model configuration
|
||
- Speaker embeddings
|
||
|
||
Those files are passed to synthesize_speech.py to generate new speech. The profile’s userId is also used when associating the generated salutation with its owner.
|
||
|
||
The synthesizer does not read training_model_s3_path; it assumes the local/EFS files remain available.
|
||
|
||
## Architectural role
|
||
|
||
User
|
||
└─ UserAudioProfile
|
||
├─ current training status
|
||
├─ current personalized-model locations
|
||
│
|
||
├── written by VoiceCloning operation
|
||
└── read by Speech Synthesis operation
|
||
|
||
It is therefore best understood as:
|
||
|
||
- A user-owned voice identity
|
||
- A readiness record
|
||
- A model-artifact manifest
|
||
- The contract joining asynchronous training and inference
|
||
|
||
It is not:
|
||
|
||
- A raw recording
|
||
- A cloning job
|
||
- A model binary
|
||
- A tier or class
|
||
- A complete description of how the model was trained
|
||
|
||
## Important weaknesses in that contract
|
||
|
||
The schema is much looser than the runtime assumptions:
|
||
|
||
- status is any string; there is no enum.
|
||
- Both model-path fields are Mixed, so their required keys are not validated.
|
||
- A profile can legally say completed while having missing or malformed paths.
|
||
- Updates do not verify that a matching profile exists.
|
||
- The two duplicated schemas can drift independently.
|
||
- Retraining appears to overwrite the same mutable profile; there is no model version history.
|
||
- S3 paths are stored but unused by the current synthesizer.
|
||
- The synthesizer’s find() returns an array and mishandles an empty result.
|
||
|
||
Overall, UserAudioProfile looks like the codebase’s canonical “current voice for this user” record: a control-plane resource pointing to whichever personalized model should presently be
|
||
used for synthesis.
|
||
|
||
18. How is S3 authenticated/authorized?
|
||
|
||
• ## Authentication
|
||
|
||
The code does not provide S3 credentials explicitly. It constructs the AWS SDK v2 client with:
|
||
|
||
const s3 = new AWS.S3()
|
||
|
||
at app/services/s3/index.js:4.
|
||
|
||
That makes the SDK use its default credential chain, including:
|
||
|
||
- AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, and optional session token
|
||
- ~/.aws/credentials
|
||
- ECS task credentials
|
||
- EC2 instance-profile credentials
|
||
|
||
The installation guide specifically instructs operators to run aws configure at Installation Guide:186 (voice-cloning/docs/potion-voice-cloning_Installation_Guide.md:186). That writes
|
||
access-key credentials to the host user’s AWS profile. This is the only documented authentication mechanism, although production could use an undocumented EC2 role.
|
||
|
||
The SDK automatically signs S3 requests using AWS Signature Version 4.
|
||
|
||
## Authorization
|
||
|
||
Authorization is entirely external to this repository. The resolved AWS identity must be permitted by IAM and the relevant bucket policies.
|
||
|
||
The application requires approximately:
|
||
|
||
- s3:PutObject for trained model uploads
|
||
- s3:PutObject for synthesized WAV uploads
|
||
- s3:GetObject if the unused fetchS3Object() helper is ever called
|
||
- Additional multipart-upload permissions when applicable
|
||
|
||
There are no IAM policies, bucket policies, Terraform files, CloudFormation templates, role definitions, or permission checks in the repository.
|
||
|
||
The code also does not set an object ACL; public-read is commented out. Object accessibility therefore depends on bucket ownership settings and bucket policies.
|
||
|
||
## What each worker accesses
|
||
|
||
- Cloning worker uploads model assets at voice-cloning-job-handler/index.js:266.
|
||
- Synthesizer worker uploads generated WAVs at voice-synthsizer-job-handler/index.js:127.
|
||
- Source recordings downloaded through CloudFront use ordinary HTTPS, not this S3 identity.
|
||
|
||
One separate concern: the cloning worker supplies potion-voice-users-training-model/${env} as the Bucket value. Normal S3 bucket names cannot contain /; the environment should likely be
|
||
part of the object key instead.
|
||
|
||
So the best-supported conclusion is: documented deployments authenticate with host-level AWS access keys created by aws configure, while all authorization is managed outside this
|
||
repository. The actual production IAM principal and permission scope cannot be determined here.
|
||
|