before changes

This commit is contained in:
2026-10-05 16:14:53 -04:00
parent dd6e6f7cfd
commit f39555ce15
94 changed files with 3285321 additions and 1 deletions

View File

@@ -0,0 +1,764 @@
# potion-voice **voice-cloning** *Installation and Usage Guide*
In this guide, you will find more detailed instructions and examples for the following tasks:
+ Setting up a new AWS GPU-backed EC2 instance suitable for training new potion-voice models;
+ Setting up software environment and (optionally) prepare data sets for training new potion-voice models;
+ Training and evaluating new potion-voice models; and
+ Usage examples for voice cloning and speech synthesizing.
## Set Up AWS GPU-backed Compute Node (non-production)
1. Set up baseline & connect to remote node:
+ GPU-enabled Compute Node (e.g., g5.2xlarge by default)
+ We recommend a GPU-enabled Compute Node with 256GB root partition (volume type: gp3; 64GB for swapfile) and 512GB secondary SDD holding all dev / data files)
+ Inbound ports: SSH and TensorBoard (e.g., port 6006)
+ Ubuntu 22.04 LTS (Server) Installation
+ SSH into the EC2 instance
1. Secure / update baseline
```sh
$ sudo apt-get update
$ sudo apt-get upgrade
$ sudo apt-get install linux-aws linux-headers-aws linux-image-aws
```
1. Disable unattended upgrades. Enter the below command and select 'No'. These Upgrades might cause version mismatch between nvidia-drivers and cuda.
```sh
$ sudo dpkg-reconfigure -plow unattended-upgrades
Replacing config file /etc/apt/apt.conf.d/20auto-upgrades with new version
```
1. Set up secondary disk (used as dev / data volume)
```sh
$ sudo lsblk
NAME MAJ:MIN RM SIZE RO TYPE MOUNTPOINT
[...]
nvme1n1 259:0 0 500G 0 disk
[...]
$ sudo mkfs -t ext4 /dev/nvme1n1
mke2fs 1.45.5 (07-Jan-2020)
Creating filesystem with 524288000 4k blocks and 131072000 inodes
Filesystem UUID: 90327770-ba4d-4003-9136-964b4388ffb6
Superblock backups stored on blocks:
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
4096000, 7962624, 11239424, 20480000, 23887872, 71663616, 78675968,
102400000, 214990848, 512000000
Allocating group tables: done
Writing inode tables: done
Creating journal (262144 blocks): done
Writing superblocks and filesystem accounting information: done
$ mkdir DEV_PATH
```
+ Edit `/etc/fstab` and add
```txt
/dev/nvme1n1 DEV_PATH ext4 defaults,nofail 0 2
```
```sh
$ sudo mount -a
$ sudo chown -R ubuntu:ubuntu DEV_PATH
$ mkdir DEV_PATH/data
```
1. Create a swap file (training is memory intensive; so, add a swap file!)
+ Use the `dd` command to create a swap file on the root file system
+ Note: The size of the swap file is the block size option multiplied by the count option in the dd command. Adjust these values to determine the desired swap file size.
+ Note: The block size you specify should be less than the available memory on the instance or you receive a "memory exhausted" error.
+ Set up the swap file (of size 64 GB [512 MB x 128]).
```sh
$ sudo dd if=/dev/zero of=/swapfile bs=512M count=128
128+0 records in
128+0 records out
68719476736 bytes (69 GB, 64 GiB) copied, 336.416 s, 204 MB/s
```
+ Update the read and write permissions for the swap file:
```sh
$ sudo chmod 600 /swapfile
```
+ Set up a Linux swap area:
```sh
$ sudo mkswap /swapfile
Setting up swapspace version 1, size = 64 GiB
no label, UUID=1dfc20ce-ed64-4e69-8fa7-a800bbea4617
```
+ Make the swap file available for immediate use by adding the swap file to swap space:
```sh
$ sudo swapon /swapfile
```
+ Verify that the procedure was successful:
```sh
$ sudo swapon -s
Filename Type Size Used Priority
/swapfile file 67108860 0 -2
```
+ Enable the swap file at boot time by editing the `/etc/fstab` file. Add the following new line at the end of the file:
```txt
/swapfile swap swap defaults 0 0
```
1. Install NVIDIA drivers / CUDA support (pytorch required version 11.6 or 12)
```sh
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.0-1_all.deb
sudo dpkg -i cuda-keyring_1.0-1_all.deb
sudo apt-get update
sudo apt-get -y install cuda-12-0
```
+ Reboot the instance and ensure all drivers load automatically
```sh
$ sudo reboot
```
+ Reconnect to the instance and verify NVIDIA drivers / CUDA support are as expected
```sh
$ nvidia-smi
Tue Jan 17 08:20:53 2023
+-----------------------------------------------------------------------------+
| NVIDIA-SMI 525.60.13 Driver Version: 525.60.13 CUDA Version: 12.0 |
|-------------------------------+----------------------+----------------------+
| GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|===============================+======================+======================|
| 0 NVIDIA A10G On | 00000000:00:1E.0 Off | 0 |
| 0% 19C P8 16W / 300W | 0MiB / 23028MiB | 0% Default |
| | | N/A |
+-------------------------------+----------------------+----------------------+
+-----------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=============================================================================|
| No running processes found |
+-----------------------------------------------------------------------------+
```
## Set Up Software Environment
1. Set up Python 3 (v3.10) development environment
```sh
$ sudo apt-get install python3-dev python3-pip python3-wheel python3-venv
```
1. Set up Phoneme back-end
```sh
$ sudo apt-get install espeak-ng espeak-ng-espeak
```
1. Set up required tools / standard dependencies
```sh
$ sudo apt-get install ffmpeg unzip git
```
1. Set up AWS Command Line Interface
```sh
$ sudo apt-get install awscli
$ aws configure
AWS Access Key ID [None]: xxxxxxxxxx
AWS Secret Access Key [None]: yyyyyyyyyy
Default region name [None]: us-west-2
Default output format [None]: json
$ aws configure set default.s3.max_concurrent_requests 50
```
1. (dev install only) Copy and extract training data sets from AWS
```sh
$ cd DEV_PATH/data
### VCTK v 0.92
$ aws s3 cp s3://potion-datasets/VCTK/VCTK-Corpus-0.92/VCTK-Corpus-0.92.tgz .
download: s3://potion-datasets/VCTK/VCTK-Corpus-0.92/VCTK-Corpus-0.92.tgz to ./VCTK-Corpus-0.92.tgz
$ tar -xzvf VCTK-Corpus-0.92.tgz
VCTK-Corpus-0.92/
VCTK-Corpus-0.92/README.txt
VCTK-Corpus-0.92/update.txt
VCTK-Corpus-0.92/license_text
VCTK-Corpus-0.92/txt/
[...]
VCTK-Corpus-0.92/wav48_silence_trimmed/p238/p238_191_mic1.flac
VCTK-Corpus-0.92/wav48_silence_trimmed/p238/p238_267_mic2.flac
$ rm VCTK-Corpus-0.92.tgz
### LibriTTS train-clean-360 subset
$ aws s3 cp s3://potion-datasets/LibriTTS/train-clean-360.tar.gz .
download: s3://potion-datasets/LibriTTS/train-clean-360.tar.gz to ./train-clean-360.tar.gz
$ tar -xzvf train-clean-360.tar.gz
./LibriTTS/train-clean-360/
./LibriTTS/train-clean-360/2272/
./LibriTTS/train-clean-360/2272/152265/
./LibriTTS/train-clean-360/2272/152265/2272_152265_000032_000001.original.txt
./LibriTTS/train-clean-360/2272/152265/2272_152265_000012_000001.wav
[...]
LibriTTS/reader_book.tsv
LibriTTS/speakers.tsv
$ rm train-clean-360.tar.gz
### Potion salutation recordings
$ aws s3 cp s3://potion-datasets/potion-voice-datasets/potion-salut-corpus_20221026.tgz .
download: s3://potion-datasets/potion-voice-datasets/potion-salut-corpus_20221026.tgz to ./potion-salut-corpus_20221019.tgz
$ tar -xzvf potion-salut-corpus_20221026.tgz
potion-salut-corpus-94de499c-b770-4e4c-97fc-6add91befe1b/
potion-salut-corpus-94de499c-b770-4e4c-97fc-6add91befe1b/wav48/
potion-salut-corpus-94de499c-b770-4e4c-97fc-6add91befe1b/wav48/POTION_6192d9c9a563df5c87ecb8bd/
potion-salut-corpus-94de499c-b770-4e4c-97fc-6add91befe1b/wav48/POTION_6192d9c9a563df5c87ecb8bd/POTION_6192d9c9a563df5c87ecb8bd_334.wav
potion-salut-corpus-94de499c-b770-4e4c-97fc-6add91befe1b/wav48/POTION_6192d9c9a563df5c87ecb8bd/POTION_6192d9c9a563df5c87ecb8bd_473.wav
[...]
potion-salut-corpus-94de499c-b770-4e4c-97fc-6add91befe1b/txt/POTION_62d82d269cbde00027b66007/POTION_62d82d269cbde00027b66007_197.txt
potion-salut-corpus-94de499c-b770-4e4c-97fc-6add91befe1b/speaker-info.txt
$ rm potion-salut-corpus_20221026.tgz
```
1. Create a virtual potion-voice-cloner working environment
```sh
$ cd DEV_PATH
$ python3 -m venv potion-voice_venv
$ cd potion-voice_venv/
$ source bin/activate
(potion-voice_venv) $
```
1. Clone the potion-voice GitHub repository
```sh
(potion-voice_venv) $ cd DEV_PATH/potion-voice_venv/
(potion-voice_venv) $ python3 -m pip install --upgrade pip
(potion-voice_venv) $ git clone https://github.com/potion/potion-voice.git
```
1. Install potion-voice requirements (dependencies) and test that PyTorch is working with the GPU properly
```sh
(potion-voice_venv) $ cd DEV_PATH/potion-voice_venv/potion-voice/
(potion-voice_venv) $ python3 -m pip install -r ./requirements.dev.txt
(potion-voice_venv) $ python3
Python 3.10.6 (main, Nov 14 2022, 16:10:14) [GCC 11.3.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> import torch
>>> torch.cuda.is_available()
True
>>> torch.cuda.get_device_name(0)
'NVIDIA A10G'
>>> quit()
```
1. Install TTS dependencies
```sh
(potion-voice_venv) $ cd voice-cloning/
(potion-voice_venv) $ git clone --depth 1 --branch v0.10.2 https://github.com/coqui-ai/TTS
(potion-voice_venv) $ python3 -m pip install -e TTS/
```
+ Note 1: Installing requirements will ask for GitHub token twice! The second request is for a dependent package, which is also a private repo.
+ Note 2: Separate requirements files have been added for development (local versus AWS) and production usage (for GPU and CPU-only deployment).
## Training New potion-voice Models (Multi-speaker Baseline & Voice Cloning)
### Preprocess Dataset(s) Required for Multi-speaker Baseline Model Training
1. For each dataset, ensure that the sampling rate matches and speaker embeddings are precomputed.
```sh
(potion-voice_venv) $ python3 prepare_datasets.py --dataset_preset vctk --dataset_archive_path ~/datasets/VCTK_v0.92/VCTK-Corpus-0.92.tgz --sampling_rate 22050
Commencing preparation of dataset for multi-speaker baseline model training:
+ Dataset preset: vctk
+ Dataset : /home/[REDACTED_HOMEDIR_USERNAME_3]/datasets/VCTK_v0.92/VCTK-Corpus-0.92.tgz
+ Output path : results/datasets
+ Sampling rate : 22050
>>> Extracting archive ...
>>> Resampling audio files to 16000Hz ...
Resampling the audio files...
Found 88328 files...
100%|████████████████████████████████████████████████████████████████████████████████| 88328/88328 [18:25<00:00, 79.88it/s]
Done !
>>> Computing speaker embeddings ...
> Found 44283 files in /home/[REDACTED_HOMEDIR_USERNAME_3]/work/potion-repos/potion-voice_venv/potion-voice/voice-cloning/results/datasets/VCTK-Corpus-0.92
> Model fully restored.
> Setting up Audio Processor...
[...]
100%|████████████████████████████████████████████████████████████████████████████████| 44283/44283 [06:18<00:00, 116.99it/s]
Speaker embeddings saved at: results/datasets/VCTK-Corpus-0.92/speakers.pth
>>> Extracting original archive again (overwritting previously resampled files)...
>>> Resampling audio files to 22050Hz ...
Resampling the audio files...
Found 88328 files...
100%|████████████████████████████████████████████████████████████████████████████████| 88328/88328 [20:48<00:00, 70.74it/s]
Done !
Completed preparing voice dataset for multi-speaker baseline model training; generated asset locations are as follows:
--> results/datasets/VCTK-Corpus-0.92
--> results/datasets/VCTK-Corpus-0.92/speakers.pth
Done; bye.
```
### Train New potion-voice Multi-speaker Baseline Model
1. To train a new baseline model:
```sh
(potion-voice_venv) $ python3 train_multispeaker_baseline_model.py
usage: train_multispeaker_baseline_model.py [-h] --datasets {VCTK,LibriTTS_tc360,POTION_Salut} [{VCTK,LibriTTS_tc360,POTION_Salut} ...] [--output_path OUTPUT_PATH] [--batch_size BATCH_SIZE] [--max_epochs MAX_EPOCHS]
Code to train multi-speaker baseline model
options:
-h, --help show this help message and exit
--datasets {VCTK,LibriTTS_tc360,POTION_Salut} [{VCTK,LibriTTS_tc360,POTION_Salut} ...]
List of training datasets to be included in training run.
--output_path OUTPUT_PATH
Path to store trained / generated assets
--batch_size BATCH_SIZE
Batch size for training run
--max_epochs MAX_EPOCHS
Maximum number of epochs for training run
```
Using the default settings, training a new multi-speaker baseline model (on an AWS g5.2xlarge instance) takes 5-7 days (100 epochs with 32 batch size and all 3 datasets (i.e., VCTK, LibriTTS_tc360, andpotion_Salut)).
1. At the end of a training run, there will be the following files in the result folder:
```txt
results/baseline-models/vits_vctk-March-23-2022_03+43AM-0000000/
|-- best_model.pth .................................... best model using avg_loss_0 (NOT the best model; suggest to ignore for now)
|-- best_model_19096.pth .............................. same as best_model.pth (suggest to ignore for now)
|-- checkpoint_300000.pth ............................. fifth last checkpoint
|-- checkpoint_310000.pth ............................. fourth last checkpoint
|-- checkpoint_320000.pth ............................. third last checkpoint
|-- checkpoint_330000.pth ............................. second last checkpoint
|-- checkpoint_340000.pth ............................. last checkpoint
|-- config.json ....................................... configuration file
|-- events.out.tfevents.1648007016.ip-172-31-83-225 ... event log for entire training run including eval samples and charts (view via tensorboard)
|-- speakers.pth ...................................... speaker embeddings
|-- trainer_0_log.txt ................................. training log
|-- train_multispeaker_baseline_model.py .............. copy of the training script
```
Use the event log to determine which of the checkpoints corresponds to the best model.
### Clone a Voice based on the Mutli-speaker Baseline Model
1. To clone a new voice, you need at least 10 voice samples (ideally 30). Those voice recordings (and their corresponding transcription files) have to be arranged as follows (and compressed into a `.tgz`, `.tbz` or `.zip` archive):
```txt
VOICE_DATASET_PATH/txt/1/1_001.txt
VOICE_DATASET_PATH/txt/1/1_002.txt
VOICE_DATASET_PATH/txt/1/1_003.txt
...
VOICE_DATASET_PATH/txt/1/1_029.txt
VOICE_DATASET_PATH/txt/1/1_030.txt
VOICE_DATASET_PATH/wav48/1/1_001.wav
VOICE_DATASET_PATH/wav48/1/1_002.wav
VOICE_DATASET_PATH/wav48/1/1_003.wav
...
VOICE_DATASET_PATH/wav48/1/1_028.wav
VOICE_DATASET_PATH/wav48/1/1_029.wav
VOICE_DATASET_PATH/wav48/1/1_030.wav
```
1. Next, pre-process audio recordings to fit the format of audio samples (i.e., sampling rate) and pre-compute speaker embeddings:
```sh
(potion-voice_venv) $ python prepare_datasets.py --dataset_preset potion_voice_cloning --dataset_archive_path ~/datasets/potion\ Recordings/potion-voice\ recordings/user123.tgz
Commencing preparation of dataset for multi-speaker baseline model training:
+ Dataset preset: potion_voice_cloning
+ Dataset : /home/[REDACTED_HOMEDIR_USERNAME_3]/datasets/potion Recordings/potion-voice recordings/user123.tgz
+ Output path : results/datasets
+ Sampling rate : 22050
>>> Extracting archive ...
>>> Resampling audio files to 16000Hz ...
Resampling the audio files...
Found 30 files...
100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 30/30 [00:00<00:00, 39.40it/s]
Done !
>>> Extracting original archive again (overwritting previously resampled files)...
>>> Resampling audio files to 22050Hz ...
Resampling the audio files...
Found 30 files...
100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 30/30 [00:00<00:00, 37.61it/s]
Done !
Completed preparing voice dataset for multi-speaker baseline model training; generated asset locations are as follows:
--> results/datasets/sr22050/user123
--> results/datasets/sr22050/user123/speakers.pth
Done; bye.
```
1. Finally, trigger voice cloning:
```sh
(potion-voice_venv) $ python3 clone_voice.py [-h] --baseline_model_path BASELINE_MODEL_PATH --speaker_dataset_path SPEAKER_DATASET_PATH --speaker_embeddings_path SPEAKER_EMBEDDINGS_PATH [--output_path OUTPUT_PATH] [--batch_size BATCH_SIZE] [--max_epochs MAX_EPOCHS] [--use_cpu] [--output_format {txt,json}]
Code to clone a voice from a given set of voice samples and a multi-speaker baseline model
options:
-h, --help show this help message and exit
--baseline_model_path BASELINE_MODEL_PATH
Path to multi-speaker baseline model (VITS model)
--speaker_dataset_path SPEAKER_DATASET_PATH
Path to voice cloning dataset
--speaker_embeddings_path SPEAKER_EMBEDDINGS_PATH
Path to speaker's embeddings file
--output_path OUTPUT_PATH
Path to store trained / generated assets
--batch_size BATCH_SIZE
Batch size for training run
--max_epochs MAX_EPOCHS
Maximum number of epochs for training run
--use_cpu Signal that CPU should be used even if a CUDA-device is available
--output_format {txt,json}
Output format; available choices include 'txt' for human readible text and 'json' for JSON formatting
```
Using the default settings and 30 audio samples, cloning a new voice (on an AWS g5.2xlarge instance) takes about one hour.
1. At the end of a voice cloning run, there will be the following files in the result folder:
```txt
results/cloned-voices/vits_potion_clone-December-28-2022_10+52AM-1327031/
|-- best_model_365097.pth .................. best model using avg_loss_0 (save to use)
|-- best_model.pth ......................... same as best_model_365097.pth
|-- checkpoint_365200.pth .................. last checkpoint
|-- clone_voice.py ......................... copy of the clone_voice script used in this run
|-- config.json ............................ configuration file
|-- events.out.tfevents.1672195961.rigel ... event log for entire voice cloning run including eval samples and charts (view via tensorboard)
|-- speakers.pth ........................... speaker's embeddings file
|-- trainer_0_log.txt ...................... training log file
```
Use the event log to confirm that the best model is indeed giving the best outputs.
### Monitoring Training Progress
Using tensorboard / tensorboardX, training progress (for both, multi-speaker baseline training and voice cloning) can be monitored and evaluation samples can be accessed.
1. Ensure AWS Security Group settings (inbound) are set appropriamust include:
```txt
HTTPS TCP 443 0.0.0.0/0
Custom_TCP TCP 6006 0.0.0.0/0
```
+ Server-side, launch the tensorboard service:
```sh
(potion-voice_venv) $ tensorboard --logdir=./results/baseline-models/vits_vctk-March-07-2022_09+47AM-0000000/ --host 0.0.0.0
TensorFlow installation not found - running with reduced feature set.
NOTE: Using experimental fast data loading logic. To disable, pass
"--load_fast=false" and report issues on GitHub. More details:
https://github.com/tensorflow/tensorboard/issues/4784
TensorBoard 2.8.0 at http://0.0.0.0:6006/ (Press CTRL+C to quit)
```
+ Locally, point your preferred Web browser to <http://PUBLIC_IPv4_DNS:6006/>
### Minimise a Cloned Voice
To minimise the size of a trained model, run the followng script which removes optimiser and discriminator components from the model -- those are only required for training but not for inference:
```sh
(potion-voice_venv) $ python3 minimize_cloned_voice_model.py [-h] --voice_model_asset_path VOICE_MODEL_ASSET_PATH [--voice_model_name VOICE_MODEL_NAME] [--voice_model_config_name VOICE_MODEL_CONFIG_NAME] [--minimise_suffix MINIMISE_SUFFIX] [--overwrite_assets] [--output_format {txt,json}]
Code to minimise (i.e., remove optimiser & discriminator) a cloned voice model
options:
-h, --help show this help message and exit
--voice_model_asset_path VOICE_MODEL_ASSET_PATH
Path to directory storing cloned voice model and the corresponding configuration and speaker files
--voice_model_name VOICE_MODEL_NAME
Name of the (best) cloned voice model
--voice_model_config_name VOICE_MODEL_CONFIG_NAME
Name of the config file for the cloned voice model
--minimise_suffix MINIMISE_SUFFIX
Suffix to be used for minimised model and its assets (i.e., new config file)
--overwrite_assets Signal whether existing model assets should be overwritten or not (default: do not overwrite)
--output_format {txt,json}
Output format; available choices include 'txt' for human readible text and 'json' for JSON formatting
```
1. Command-line output sample for output format option "txt":
```sh
$ python3 minimize_cloned_voice_model.py --voice_model_asset_path results/cloned-voices/vits_potion_clone-December-28-2022_10+52AM-1327031/
Minimising given voice model:
+ Cloned voice model file path : results/cloned-voices/vits_potion_clone-December-28-2022_10+52AM-1327031/best_model.pth
+ Cloned voice model config file : results/cloned-voices/vits_potion_clone-December-28-2022_10+52AM-1327031/config.json
> Using model: vits
> Setting up Audio Processor...
[...]
Completed minimising cloned voice model. The resulting (modified) assets can be found at:
--> Minimised voice model path : results/cloned-voices/vits_potion_clone-December-28-2022_10+52AM-1327031/best_model_light.pth
--> Minimised voice model config path: results/cloned-voices/vits_potion_clone-December-28-2022_10+52AM-1327031/config_light.json
Done; bye.
```
### Scoring a Cloned Voice
1. To score a cloned voice, run the following command:
```sh
(potion-voice_venv) $ python3 score_cloned_voice.py [-h] --voice_dataset_path VOICE_DATASET_PATH --voice_model_path VOICE_MODEL_PATH --voice_model_config_path VOICE_MODEL_CONFIG_PATH --speaker_embeddings_path SPEAKER_EMBEDDINGS_PATH [--temp_path TEMP_PATH] [--keep_temp] [--use_cpu] [--output_format {txt,json}]
Compute quality score for a given voice model (cloned voice) wrt. a given set of voice recordings (original voice))
options:
-h, --help show this help message and exit
--voice_dataset_path VOICE_DATASET_PATH
Path to set of voice recordings (original voice)
--voice_model_path VOICE_MODEL_PATH
Path to cloned voice model
--voice_model_config_path VOICE_MODEL_CONFIG_PATH
Path to config file for the cloned voice model
--speaker_embeddings_path SPEAKER_EMBEDDINGS_PATH
Path to speaker's embeddings file (i.e., pre-computed embeddings typically stored with the speaker's dataset)
--temp_path TEMP_PATH
Path to store temporary speech assets
--keep_temp Signal that temporary assets used for scoring should not be deleted once done
--use_cpu Signal that CPU should be used even if a CUDA-device is available
--output_format {txt,json}
Output format; available choices include 'txt' for human readible text and 'json' for JSON formatting
```
1. Command-line output sample for output format option "txt":
```sh
$ python3 score_cloned_voice.py --voice_dataset_path results/datasets/sr22050/michael/wav48/1/ --voice_model_path results/cloned-voices/vits_potion_clone-December-28-2022_01+09AM-1327031/best_model.pth --voice_model_config_path results/cloned-voices/vits_potion_clone-December-28-2022_01+09AM-1327031/config.json --speaker_embeddings_path results/datasets/sr22050/michael/speakers.pth
Computing similarity score for a given voice model (cloned voice) wrt. a given set of voice recordings (original voice):
+ Original voice recordings path: results/datasets/sr22050/michael/wav48/1/
+ Cloned voice model file path : results/cloned-voices/vits_potion_clone-December-28-2022_01+09AM-1327031/best_model.pth
+ Cloned voice model config file: results/cloned-voices/vits_potion_clone-December-28-2022_01+09AM-1327031/config.json
+ Speaker embeddings file : results/datasets/sr22050/michael/speakers.pth
+ CUDA availability : True
+ Compute device used : cuda
+ No. of speakers : 1
+ Speaker's names : ['VCTK_old_1']
+ No. of embeddings : 30
> Using model: vits
> Setting up Audio Processor...
Loaded the voice encoder model on cuda in 0.01 seconds.
Completed computing similarity score for the two sets of recordings. The resulting similarity score is:
--> 0.9127510190010071
Done; bye.
```
1. Command-line output sample for output format option "json":
```sh
(potion-voice_venv)$ python3 score_cloned_voice.py --voice_dataset_path results/datasets/sr22050/michael/wav48/1/ --voice_model_path results/cloned-voices/vits_potion_clone-December-28-2022_01+09AM-1327031/best_model.pth --voice_model_config_path results/cloned-voices/vits_potion_clone-December-28-2022_01+09AM-1327031/config.json --speaker_embeddings_path results/datasets/sr22050/michael/speakers.pth --output_format json
> Using model: vits
> Setting up Audio Processor...
[...]
Loaded the voice encoder model on cuda in 0.01 seconds.
{"success": true, "in": {"voice_dataset_path": "results/datasets/sr22050/michael/wav48/1/", "voice_model_path": "results/cloned-voices/vits_potion_clone-December-28-2022_01+09AM-1327031/best_model.pth"}, "out": {"score": 0.91}}
```
## Usage Examples for Speech Synthesizing
1. To generate speech for a given cloned voice, run the following command:
```sh
(potion-voice_venv) $ python3 synthesize_speech.py [-h] --voice_model_path VOICE_MODEL_PATH --voice_model_config_path VOICE_MODEL_CONFIG_PATH --speaker_embeddings_path SPEAKER_EMBEDDINGS_PATH --txt TXT [--output_path OUTPUT_PATH] [--target_sampling_rate TARGET_SAMPLING_RATE] [--speech_sample_wav_path SPEECH_SAMPLE_WAV_PATH] [--speech_sample_txt SPEECH_SAMPLE_TXT] [--trim_silence] [--use_cpu] [--output_format {txt,json}]
Code to synthesize speech for a given voice model
options:
-h, --help show this help message and exit
--voice_model_path VOICE_MODEL_PATH
Path to cloned voice model
--voice_model_config_path VOICE_MODEL_CONFIG_PATH
Path to config file for the cloned voice model
--speaker_embeddings_path SPEAKER_EMBEDDINGS_PATH
Path to speaker's embeddings file (i.e., pre-computed embeddings typically stored with the speaker's dataset)
--txt TXT Text to synthesize
--output_path OUTPUT_PATH
Path to store generated speech assets
--target_sampling_rate TARGET_SAMPLING_RATE
Desired sampling rate (in Hz) for output file
--speech_sample_wav_path SPEECH_SAMPLE_WAV_PATH
Path to a sample utterance of the speaker (used for style transfer)
--speech_sample_txt SPEECH_SAMPLE_TXT
Text of the sample utterance of the speaker (used for style transfer)
--trim_silence Signal whether to trim silence from synthesised speech
--use_cpu Signal that CPU should be used even if a CUDA-device is available
--output_format {txt,json}
Output format; available choices include 'txt' for human readible text and 'json' for JSON formatting
```
1. Command-line output sample for output format option "txt":
```sh
(potion-voice_venv) $ python3 synthesize_speech.py --voice_model_path results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/best_model.pth --voice_model_config_path results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/config.json --speaker_embeddings_path results/datasets/sr22050/[REDACTED_HOMEDIR_USERNAME_2]/speakers.pth --txt "Hi person_82, it works!"
Commencing speech synthesizing:
+ Voice model file path : results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/best_model.pth
+ Voice model config file: results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/config.json
+ Speaker embeddings file: results/datasets/sr22050/[REDACTED_HOMEDIR_USERNAME_2]/speakers.pth
+ Output path : results/speech
+ Text to synthesize : Hi person_82, it works!
+ CUDA availability : True
+ Compute device used : cuda
+ No. of speakers : 1
+ Speaker's names : ['VCTK_old_1']
+ No. of embeddings : 30
> Using model: vits
> Setting up Audio Processor...
[...]
>>> Saving original output to : results/speech/b4189e9e-6142-4dad-8577-6de77087ffd1.wav
>>> Saving resampled output to: results/speech/b4189e9e-6142-4dad-8577-6de77087ffd1_sr48000.wav
Speech synthesizing has completed. Bye.
```
1. Command-line output sample for output format option "txt":
```sh
(potion-voice_venv) $ python3 synthesize_speech.py --voice_model_path results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/best_model_light.pth --voice_model_config_path results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/config_light.json --speaker_embeddings_path results/datasets/sr22050/[REDACTED_HOMEDIR_USERNAME_2]/speakers.pth --txt "Hi person_82, it works!" --output_format json
> Using model: vits
> Setting up Audio Processor...
[...]
{"success": true, "in": {"voice_model_path": "results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/best_model_light.pth", "voice_model_config_path": "results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/config_light.json", "speaker_embeddings_path": "results/datasets/sr22050/[REDACTED_HOMEDIR_USERNAME_2]/speakers.pth"}, "out": {"speech_original_path": "results/speech/c99e494c-e1f9-4c12-9095-255cf7db792b.wav", "speech_resampled_path": "results/speech/c99e494c-e1f9-4c12-9095-255cf7db792b_sr48000.wav"}}
```
### Scoring a Synthesised Salutation
1. To score a synthesised salutation, run the following command:
```sh
(potion-voice_venv)$ python3 score_salutation.py [-h] --recording_path RECORDING_PATH --first_name FIRST_NAME [--output_format {txt,json}]
Score a given salutation recording wrt. its desired content, the actual salutation recording, and a generated transcription (using Potion's internal Transciption API) of the recording.
optional arguments:
-h, --help show this help message and exit
--recording_path RECORDING_PATH
Path to salutation recoding (.wav audio file)
--first_name FIRST_NAME
First name that the salutation recoding is meant to use
--output_format {txt,json}
Output format; available choices include 'txt' for human readible text and 'json' for JSON formatting
```
1. Command-line output sample for output format option "txt":
```sh
(potion-voice_venv)$ python3 score_salutation.py --recording_path /home/[REDACTED_HOMEDIR_USERNAME_3]/person_82_-_Hey_person_83.wav --first_name person_83
Commencing scoring of the given salutation recording:
+ Salutation recording path: /home/[REDACTED_HOMEDIR_USERNAME_3]/person_82_-_Hey_person_83.wav
+ Salutation first name : person_83
>> Salutation score : 0.892155
Done; bye.
```
1. Command-line output sample for output format option "json":
```sh
(potion-voice_venv)$ python3 score_salutation.py --recording_path /home/[REDACTED_HOMEDIR_USERNAME_3]/person_82_-_Hey_person_83.wav --first_name person_83 --output_format json
{"in": {"recording_path": "/home/[REDACTED_HOMEDIR_USERNAME_3]/person_82_-_Hey_person_83.wav", "first_name": "person_83"}, "out": {"score": 0.89}}
```
## Troubleshooting
1. How to better monitor GPU load / utilisation?
+ Install an interactive NVIDIA-GPU process viewer such as `nvitop`:
```sh
$ python3 -m pip install nvitop
Collecting nvitop
[...]
Installing collected packages: nvidia-ml-py, termcolor, psutil, nvitop
Successfully installed nvidia-ml-py-11.495.46 nvitop-0.8.0 psutil-5.9.2 termcolor-2.0.1
````
+ Run via command-line: `nvitop`:
```sh
Tue Sep 13 02:09:07 2022
╒═════════════════════════════════════════════════════════════════════════════╕
│ NVITOP 0.8.0 Driver Version: 515.65.01 CUDA Driver Version: 11.7 │
├───────────────────────────────┬──────────────────────┬──────────────────────┤
│ GPU Name Persistence-M│ Bus-Id Disp.A │ Volatile Uncorr. ECC │
│ Fan Temp Perf Pwr:Usage/Cap│ Memory-Usage │ GPU-Util Compute M. │
╞═══════════════════════════════╪══════════════════════╪══════════════════════╪══════════════════════════╕
│ 0 A10G On │ 00000000:00:1E.0 Off │ 0 │ MEM: █████████▊ 65.1% │
│ 0% 47C P0 192W / 300W │ 14982MiB / 22.49GiB │ 100% Default │ UTL: ███████████████ MAX │
╘═══════════════════════════════╧══════════════════════╧══════════════════════╧══════════════════════════╛
[ CPU: ██████████▏ 18.1% ] ( Load Average: 1.07 1.11 1.04 )
[ MEM: ███████████▎ 20.2% ] [ SWP: ▏ 0.3% ]
╒════════════════════════════════════════════════════════════════════════════════════════════════════════╕
│ Processes: ubuntu@ip-172-31-95-84 │
│ GPU PID USER GPU-MEM %SM %CPU %MEM TIME COMMAND │
╞════════════════════════════════════════════════════════════════════════════════════════════════════════╡
│ 0 2100 C ubuntu 14463MiB 90 103.7 9.6 5.4 days python3 train_multispeaker_baseline_model.py │
╘════════════════════════════════════════════════════════════════════════════════════════════════════════╛
```