lots of change - all to start my 3rd redo
This commit is contained in:
118
sources/git-arch-sources/260911C-pro-v2.md
Normal file
118
sources/git-arch-sources/260911C-pro-v2.md
Normal file
@@ -0,0 +1,118 @@
|
||||
|
||||
• The pro_v2 implementation is complete but currently uncommitted and unstaged. The temporary commit was removed; HEAD is back at 26ba3d1.
|
||||
|
||||
## Problem addressed
|
||||
|
||||
The worker previously had no tier concept:
|
||||
|
||||
- Queue validation ignored pro_v2.
|
||||
- Mongoose schemas could not durably represent the tier.
|
||||
- Completion detection considered only status and asset presence.
|
||||
- A pro_v2 request could therefore be acknowledged using completed legacy assets without running the requested tier.
|
||||
- Local and uploaded artifacts had no tier namespace, allowing cross-tier reuse.
|
||||
|
||||
## Tier contract
|
||||
|
||||
A new centralized tier module was added in voice-cloning-job-handler/cloning_tiers.js:1.
|
||||
|
||||
It:
|
||||
|
||||
- Defines pro_v2 as the supported tier.
|
||||
- Treats an omitted or null tier as the existing legacy behavior.
|
||||
- Accepts tier information from:
|
||||
- tier
|
||||
- _doc.tier
|
||||
- _doc.metadata.tier
|
||||
|
||||
- Normalizes accepted values into _doc.tier.
|
||||
- Rejects blank, whitespace-padded, conflicting, or unsupported tier values.
|
||||
- Reads fields from both ordinary objects and Mongoose _doc objects.
|
||||
- Provides common comparison helpers for jobs, cloning records, and audio profiles.
|
||||
|
||||
## Queue processing changes
|
||||
|
||||
voice-cloning-job-handler/queue_worker.js:36 now validates and normalizes the tier with the rest of the queue payload.
|
||||
|
||||
After loading MongoDB state, the worker:
|
||||
|
||||
1. Resolves the tier from the message and stored cloning record.
|
||||
2. Rejects a request if both contain different non-null tiers.
|
||||
3. Falls back to the stored tier during redelivery if the message does not contain one.
|
||||
4. Passes the normalized tier into the training pipeline.
|
||||
|
||||
Completion detection is now tier-aware. A job counts as already completed only when:
|
||||
|
||||
- Both records are completed.
|
||||
- Both local and S3 asset maps are complete.
|
||||
- The VoiceCloning.tier matches the requested tier.
|
||||
- The profile’s training_model_tier matches the requested tier.
|
||||
|
||||
Consequently, completed legacy assets cannot short-circuit a new pro_v2 request.
|
||||
|
||||
During processing, the worker persists the tier on the cloning record. After training, it atomically associates the returned asset maps with training_model_tier on the profile. It
|
||||
verifies the returned Mongo documents contain the expected status, assets, and tier before recording the final cloning completion state and acknowledging SQS.
|
||||
|
||||
The existing visibility heartbeat, retry backoff, and delayed acknowledgement behavior remains unchanged.
|
||||
|
||||
## Artifact isolation
|
||||
|
||||
voice-cloning-job-handler/training_pipeline.js:240 now namespaces tiered artifacts.
|
||||
|
||||
Legacy paths remain unchanged:
|
||||
|
||||
/tmp/<directoryName>
|
||||
<efsRoot>/<env>/<directoryName>
|
||||
<directoryName>/<asset> # S3 object key
|
||||
|
||||
pro_v2 paths become:
|
||||
|
||||
/tmp/pro_v2/<directoryName>
|
||||
<efsRoot>/<env>/pro_v2/<directoryName>
|
||||
pro_v2/<directoryName>/<asset> # S3 object key
|
||||
|
||||
The archive command now runs from the tier-specific temporary root.
|
||||
|
||||
Existing local or S3 assets are reused only if the profile’s stored model tier matches the requested tier. Crash recovery still discovers completed EFS assets, but only inside the
|
||||
appropriate tier namespace.
|
||||
|
||||
The S3 bucket selection was not changed—only the object-key prefix was made tier-aware.
|
||||
|
||||
## Database schemas
|
||||
|
||||
The following optional fields were added:
|
||||
|
||||
- VoiceCloning.tier
|
||||
- UserAudioProfile.training_model_tier
|
||||
|
||||
They accept null for legacy records or pro_v2 for tiered records.
|
||||
|
||||
The duplicated schemas were all updated so the fields are not stripped depending on which worker or service loads them:
|
||||
|
||||
- Handler VoiceCloning model (voice-cloning-job-handler/voice_cloning/voice_cloning_model.js:18)
|
||||
- Shared VoiceCloning model (app/services/voice_cloning/voice_cloning_model.js:18)
|
||||
- Cloning profile model (voice-cloning-job-handler/user_audio_profile/user_audio_profile_model.js:26)
|
||||
- Synthesis profile model (voice-synthsizer-job-handler/user_audio_profile/user_audio_profile_model.js:26)
|
||||
|
||||
MongoDB is schemaless and these fields are optional, so no migration is required.
|
||||
|
||||
## Tests added
|
||||
|
||||
The queue tests now verify that:
|
||||
|
||||
- A pro_v2 request does not reuse or complete from legacy assets.
|
||||
- The pipeline receives the normalized pro_v2 tier.
|
||||
- Successful processing persists tier state on both records.
|
||||
- A subsequent matching pro_v2 redelivery skips duplicate training.
|
||||
- Already-completed matching pro_v2 state remains idempotent.
|
||||
- All three supported queue-field locations are normalized.
|
||||
- Unsupported tiers are rejected.
|
||||
- Both legacy-null and pro_v2 Mongoose documents validate.
|
||||
|
||||
All 28 tests pass, JavaScript syntax checks pass, and git diff --check reports no formatting errors.
|
||||
|
||||
## Important scope detail
|
||||
|
||||
pro_v2 currently runs the existing VITS training sequence and checkpoints. This change provides correct routing, state tracking, retries, and artifact isolation; it does not introduce
|
||||
a separate Python model, checkpoint, or hyperparameter set for pro_v2, because none exists in this repository.
|
||||
|
||||
The behavior is documented in README.md:23.
|
||||
Reference in New Issue
Block a user