before changes
This commit is contained in:
@@ -0,0 +1,764 @@
|
||||
# potion-voice **voice-cloning** *Installation and Usage Guide*
|
||||
|
||||
In this guide, you will find more detailed instructions and examples for the following tasks:
|
||||
|
||||
+ Setting up a new AWS GPU-backed EC2 instance suitable for training new potion-voice models;
|
||||
+ Setting up software environment and (optionally) prepare data sets for training new potion-voice models;
|
||||
+ Training and evaluating new potion-voice models; and
|
||||
+ Usage examples for voice cloning and speech synthesizing.
|
||||
|
||||
## Set Up AWS GPU-backed Compute Node (non-production)
|
||||
|
||||
1. Set up baseline & connect to remote node:
|
||||
|
||||
+ GPU-enabled Compute Node (e.g., g5.2xlarge by default)
|
||||
+ We recommend a GPU-enabled Compute Node with 256GB root partition (volume type: gp3; 64GB for swapfile) and 512GB secondary SDD holding all dev / data files)
|
||||
+ Inbound ports: SSH and TensorBoard (e.g., port 6006)
|
||||
+ Ubuntu 22.04 LTS (Server) Installation
|
||||
+ SSH into the EC2 instance
|
||||
|
||||
1. Secure / update baseline
|
||||
|
||||
```sh
|
||||
$ sudo apt-get update
|
||||
$ sudo apt-get upgrade
|
||||
$ sudo apt-get install linux-aws linux-headers-aws linux-image-aws
|
||||
```
|
||||
|
||||
1. Disable unattended upgrades. Enter the below command and select 'No'. These Upgrades might cause version mismatch between nvidia-drivers and cuda.
|
||||
|
||||
```sh
|
||||
$ sudo dpkg-reconfigure -plow unattended-upgrades
|
||||
Replacing config file /etc/apt/apt.conf.d/20auto-upgrades with new version
|
||||
```
|
||||
|
||||
1. Set up secondary disk (used as dev / data volume)
|
||||
|
||||
```sh
|
||||
$ sudo lsblk
|
||||
|
||||
NAME MAJ:MIN RM SIZE RO TYPE MOUNTPOINT
|
||||
[...]
|
||||
nvme1n1 259:0 0 500G 0 disk
|
||||
[...]
|
||||
|
||||
$ sudo mkfs -t ext4 /dev/nvme1n1
|
||||
|
||||
mke2fs 1.45.5 (07-Jan-2020)
|
||||
Creating filesystem with 524288000 4k blocks and 131072000 inodes
|
||||
Filesystem UUID: 90327770-ba4d-4003-9136-964b4388ffb6
|
||||
Superblock backups stored on blocks:
|
||||
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||
4096000, 7962624, 11239424, 20480000, 23887872, 71663616, 78675968,
|
||||
102400000, 214990848, 512000000
|
||||
|
||||
Allocating group tables: done
|
||||
Writing inode tables: done
|
||||
Creating journal (262144 blocks): done
|
||||
Writing superblocks and filesystem accounting information: done
|
||||
|
||||
$ mkdir DEV_PATH
|
||||
```
|
||||
|
||||
+ Edit `/etc/fstab` and add
|
||||
|
||||
```txt
|
||||
/dev/nvme1n1 DEV_PATH ext4 defaults,nofail 0 2
|
||||
```
|
||||
|
||||
```sh
|
||||
$ sudo mount -a
|
||||
$ sudo chown -R ubuntu:ubuntu DEV_PATH
|
||||
$ mkdir DEV_PATH/data
|
||||
```
|
||||
|
||||
1. Create a swap file (training is memory intensive; so, add a swap file!)
|
||||
|
||||
+ Use the `dd` command to create a swap file on the root file system
|
||||
+ Note: The size of the swap file is the block size option multiplied by the count option in the dd command. Adjust these values to determine the desired swap file size.
|
||||
+ Note: The block size you specify should be less than the available memory on the instance or you receive a "memory exhausted" error.
|
||||
|
||||
+ Set up the swap file (of size 64 GB [512 MB x 128]).
|
||||
|
||||
```sh
|
||||
$ sudo dd if=/dev/zero of=/swapfile bs=512M count=128
|
||||
128+0 records in
|
||||
128+0 records out
|
||||
68719476736 bytes (69 GB, 64 GiB) copied, 336.416 s, 204 MB/s
|
||||
```
|
||||
|
||||
+ Update the read and write permissions for the swap file:
|
||||
|
||||
```sh
|
||||
$ sudo chmod 600 /swapfile
|
||||
```
|
||||
|
||||
+ Set up a Linux swap area:
|
||||
|
||||
```sh
|
||||
$ sudo mkswap /swapfile
|
||||
Setting up swapspace version 1, size = 64 GiB
|
||||
no label, UUID=1dfc20ce-ed64-4e69-8fa7-a800bbea4617
|
||||
```
|
||||
|
||||
+ Make the swap file available for immediate use by adding the swap file to swap space:
|
||||
|
||||
```sh
|
||||
$ sudo swapon /swapfile
|
||||
```
|
||||
|
||||
+ Verify that the procedure was successful:
|
||||
|
||||
```sh
|
||||
$ sudo swapon -s
|
||||
Filename Type Size Used Priority
|
||||
/swapfile file 67108860 0 -2
|
||||
```
|
||||
|
||||
+ Enable the swap file at boot time by editing the `/etc/fstab` file. Add the following new line at the end of the file:
|
||||
|
||||
```txt
|
||||
/swapfile swap swap defaults 0 0
|
||||
```
|
||||
|
||||
1. Install NVIDIA drivers / CUDA support (pytorch required version 11.6 or 12)
|
||||
|
||||
```sh
|
||||
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.0-1_all.deb
|
||||
sudo dpkg -i cuda-keyring_1.0-1_all.deb
|
||||
sudo apt-get update
|
||||
sudo apt-get -y install cuda-12-0
|
||||
```
|
||||
|
||||
+ Reboot the instance and ensure all drivers load automatically
|
||||
|
||||
```sh
|
||||
$ sudo reboot
|
||||
```
|
||||
|
||||
+ Reconnect to the instance and verify NVIDIA drivers / CUDA support are as expected
|
||||
|
||||
```sh
|
||||
$ nvidia-smi
|
||||
|
||||
Tue Jan 17 08:20:53 2023
|
||||
+-----------------------------------------------------------------------------+
|
||||
| NVIDIA-SMI 525.60.13 Driver Version: 525.60.13 CUDA Version: 12.0 |
|
||||
|-------------------------------+----------------------+----------------------+
|
||||
| GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC |
|
||||
| Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. |
|
||||
| | | MIG M. |
|
||||
|===============================+======================+======================|
|
||||
| 0 NVIDIA A10G On | 00000000:00:1E.0 Off | 0 |
|
||||
| 0% 19C P8 16W / 300W | 0MiB / 23028MiB | 0% Default |
|
||||
| | | N/A |
|
||||
+-------------------------------+----------------------+----------------------+
|
||||
|
||||
+-----------------------------------------------------------------------------+
|
||||
| Processes: |
|
||||
| GPU GI CI PID Type Process name GPU Memory |
|
||||
| ID ID Usage |
|
||||
|=============================================================================|
|
||||
| No running processes found |
|
||||
+-----------------------------------------------------------------------------+
|
||||
```
|
||||
|
||||
## Set Up Software Environment
|
||||
|
||||
1. Set up Python 3 (v3.10) development environment
|
||||
|
||||
```sh
|
||||
$ sudo apt-get install python3-dev python3-pip python3-wheel python3-venv
|
||||
```
|
||||
|
||||
1. Set up Phoneme back-end
|
||||
|
||||
```sh
|
||||
$ sudo apt-get install espeak-ng espeak-ng-espeak
|
||||
```
|
||||
|
||||
1. Set up required tools / standard dependencies
|
||||
|
||||
```sh
|
||||
$ sudo apt-get install ffmpeg unzip git
|
||||
```
|
||||
|
||||
1. Set up AWS Command Line Interface
|
||||
|
||||
```sh
|
||||
$ sudo apt-get install awscli
|
||||
$ aws configure
|
||||
|
||||
AWS Access Key ID [None]: xxxxxxxxxx
|
||||
AWS Secret Access Key [None]: yyyyyyyyyy
|
||||
Default region name [None]: us-west-2
|
||||
Default output format [None]: json
|
||||
|
||||
$ aws configure set default.s3.max_concurrent_requests 50
|
||||
```
|
||||
|
||||
1. (dev install only) Copy and extract training data sets from AWS
|
||||
|
||||
```sh
|
||||
$ cd DEV_PATH/data
|
||||
|
||||
### VCTK v 0.92
|
||||
$ aws s3 cp s3://potion-datasets/VCTK/VCTK-Corpus-0.92/VCTK-Corpus-0.92.tgz .
|
||||
download: s3://potion-datasets/VCTK/VCTK-Corpus-0.92/VCTK-Corpus-0.92.tgz to ./VCTK-Corpus-0.92.tgz
|
||||
|
||||
$ tar -xzvf VCTK-Corpus-0.92.tgz
|
||||
VCTK-Corpus-0.92/
|
||||
VCTK-Corpus-0.92/README.txt
|
||||
VCTK-Corpus-0.92/update.txt
|
||||
VCTK-Corpus-0.92/license_text
|
||||
VCTK-Corpus-0.92/txt/
|
||||
[...]
|
||||
VCTK-Corpus-0.92/wav48_silence_trimmed/p238/p238_191_mic1.flac
|
||||
VCTK-Corpus-0.92/wav48_silence_trimmed/p238/p238_267_mic2.flac
|
||||
|
||||
$ rm VCTK-Corpus-0.92.tgz
|
||||
|
||||
### LibriTTS train-clean-360 subset
|
||||
$ aws s3 cp s3://potion-datasets/LibriTTS/train-clean-360.tar.gz .
|
||||
download: s3://potion-datasets/LibriTTS/train-clean-360.tar.gz to ./train-clean-360.tar.gz
|
||||
|
||||
$ tar -xzvf train-clean-360.tar.gz
|
||||
./LibriTTS/train-clean-360/
|
||||
./LibriTTS/train-clean-360/2272/
|
||||
./LibriTTS/train-clean-360/2272/152265/
|
||||
./LibriTTS/train-clean-360/2272/152265/2272_152265_000032_000001.original.txt
|
||||
./LibriTTS/train-clean-360/2272/152265/2272_152265_000012_000001.wav
|
||||
[...]
|
||||
LibriTTS/reader_book.tsv
|
||||
LibriTTS/speakers.tsv
|
||||
|
||||
$ rm train-clean-360.tar.gz
|
||||
|
||||
### Potion salutation recordings
|
||||
$ aws s3 cp s3://potion-datasets/potion-voice-datasets/potion-salut-corpus_20221026.tgz .
|
||||
download: s3://potion-datasets/potion-voice-datasets/potion-salut-corpus_20221026.tgz to ./potion-salut-corpus_20221019.tgz
|
||||
|
||||
$ tar -xzvf potion-salut-corpus_20221026.tgz
|
||||
potion-salut-corpus-94de499c-b770-4e4c-97fc-6add91befe1b/
|
||||
potion-salut-corpus-94de499c-b770-4e4c-97fc-6add91befe1b/wav48/
|
||||
potion-salut-corpus-94de499c-b770-4e4c-97fc-6add91befe1b/wav48/POTION_6192d9c9a563df5c87ecb8bd/
|
||||
potion-salut-corpus-94de499c-b770-4e4c-97fc-6add91befe1b/wav48/POTION_6192d9c9a563df5c87ecb8bd/POTION_6192d9c9a563df5c87ecb8bd_334.wav
|
||||
potion-salut-corpus-94de499c-b770-4e4c-97fc-6add91befe1b/wav48/POTION_6192d9c9a563df5c87ecb8bd/POTION_6192d9c9a563df5c87ecb8bd_473.wav
|
||||
[...]
|
||||
potion-salut-corpus-94de499c-b770-4e4c-97fc-6add91befe1b/txt/POTION_62d82d269cbde00027b66007/POTION_62d82d269cbde00027b66007_197.txt
|
||||
potion-salut-corpus-94de499c-b770-4e4c-97fc-6add91befe1b/speaker-info.txt
|
||||
|
||||
$ rm potion-salut-corpus_20221026.tgz
|
||||
```
|
||||
|
||||
1. Create a virtual potion-voice-cloner working environment
|
||||
|
||||
```sh
|
||||
$ cd DEV_PATH
|
||||
$ python3 -m venv potion-voice_venv
|
||||
$ cd potion-voice_venv/
|
||||
$ source bin/activate
|
||||
(potion-voice_venv) $
|
||||
```
|
||||
|
||||
1. Clone the potion-voice GitHub repository
|
||||
|
||||
```sh
|
||||
(potion-voice_venv) $ cd DEV_PATH/potion-voice_venv/
|
||||
(potion-voice_venv) $ python3 -m pip install --upgrade pip
|
||||
(potion-voice_venv) $ git clone https://github.com/potion/potion-voice.git
|
||||
```
|
||||
|
||||
1. Install potion-voice requirements (dependencies) and test that PyTorch is working with the GPU properly
|
||||
|
||||
```sh
|
||||
(potion-voice_venv) $ cd DEV_PATH/potion-voice_venv/potion-voice/
|
||||
(potion-voice_venv) $ python3 -m pip install -r ./requirements.dev.txt
|
||||
(potion-voice_venv) $ python3
|
||||
Python 3.10.6 (main, Nov 14 2022, 16:10:14) [GCC 11.3.0] on linux
|
||||
Type "help", "copyright", "credits" or "license" for more information.
|
||||
>>> import torch
|
||||
>>> torch.cuda.is_available()
|
||||
True
|
||||
>>> torch.cuda.get_device_name(0)
|
||||
'NVIDIA A10G'
|
||||
>>> quit()
|
||||
```
|
||||
|
||||
1. Install TTS dependencies
|
||||
|
||||
```sh
|
||||
(potion-voice_venv) $ cd voice-cloning/
|
||||
(potion-voice_venv) $ git clone --depth 1 --branch v0.10.2 https://github.com/coqui-ai/TTS
|
||||
(potion-voice_venv) $ python3 -m pip install -e TTS/
|
||||
```
|
||||
|
||||
+ Note 1: Installing requirements will ask for GitHub token twice! The second request is for a dependent package, which is also a private repo.
|
||||
|
||||
+ Note 2: Separate requirements files have been added for development (local versus AWS) and production usage (for GPU and CPU-only deployment).
|
||||
|
||||
## Training New potion-voice Models (Multi-speaker Baseline & Voice Cloning)
|
||||
|
||||
### Preprocess Dataset(s) Required for Multi-speaker Baseline Model Training
|
||||
|
||||
1. For each dataset, ensure that the sampling rate matches and speaker embeddings are precomputed.
|
||||
|
||||
```sh
|
||||
(potion-voice_venv) $ python3 prepare_datasets.py --dataset_preset vctk --dataset_archive_path ~/datasets/VCTK_v0.92/VCTK-Corpus-0.92.tgz --sampling_rate 22050
|
||||
Commencing preparation of dataset for multi-speaker baseline model training:
|
||||
|
||||
+ Dataset preset: vctk
|
||||
+ Dataset : /home/[REDACTED_HOMEDIR_USERNAME_3]/datasets/VCTK_v0.92/VCTK-Corpus-0.92.tgz
|
||||
+ Output path : results/datasets
|
||||
+ Sampling rate : 22050
|
||||
|
||||
>>> Extracting archive ...
|
||||
>>> Resampling audio files to 16000Hz ...
|
||||
Resampling the audio files...
|
||||
Found 88328 files...
|
||||
100%|████████████████████████████████████████████████████████████████████████████████| 88328/88328 [18:25<00:00, 79.88it/s]
|
||||
Done !
|
||||
>>> Computing speaker embeddings ...
|
||||
> Found 44283 files in /home/[REDACTED_HOMEDIR_USERNAME_3]/work/potion-repos/potion-voice_venv/potion-voice/voice-cloning/results/datasets/VCTK-Corpus-0.92
|
||||
> Model fully restored.
|
||||
> Setting up Audio Processor...
|
||||
[...]
|
||||
100%|████████████████████████████████████████████████████████████████████████████████| 44283/44283 [06:18<00:00, 116.99it/s]
|
||||
Speaker embeddings saved at: results/datasets/VCTK-Corpus-0.92/speakers.pth
|
||||
>>> Extracting original archive again (overwritting previously resampled files)...
|
||||
>>> Resampling audio files to 22050Hz ...
|
||||
Resampling the audio files...
|
||||
Found 88328 files...
|
||||
100%|████████████████████████████████████████████████████████████████████████████████| 88328/88328 [20:48<00:00, 70.74it/s]
|
||||
Done !
|
||||
|
||||
Completed preparing voice dataset for multi-speaker baseline model training; generated asset locations are as follows:
|
||||
--> results/datasets/VCTK-Corpus-0.92
|
||||
--> results/datasets/VCTK-Corpus-0.92/speakers.pth
|
||||
|
||||
Done; bye.
|
||||
```
|
||||
|
||||
### Train New potion-voice Multi-speaker Baseline Model
|
||||
|
||||
1. To train a new baseline model:
|
||||
|
||||
```sh
|
||||
(potion-voice_venv) $ python3 train_multispeaker_baseline_model.py
|
||||
|
||||
usage: train_multispeaker_baseline_model.py [-h] --datasets {VCTK,LibriTTS_tc360,POTION_Salut} [{VCTK,LibriTTS_tc360,POTION_Salut} ...] [--output_path OUTPUT_PATH] [--batch_size BATCH_SIZE] [--max_epochs MAX_EPOCHS]
|
||||
|
||||
Code to train multi-speaker baseline model
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--datasets {VCTK,LibriTTS_tc360,POTION_Salut} [{VCTK,LibriTTS_tc360,POTION_Salut} ...]
|
||||
List of training datasets to be included in training run.
|
||||
--output_path OUTPUT_PATH
|
||||
Path to store trained / generated assets
|
||||
--batch_size BATCH_SIZE
|
||||
Batch size for training run
|
||||
--max_epochs MAX_EPOCHS
|
||||
Maximum number of epochs for training run
|
||||
```
|
||||
|
||||
Using the default settings, training a new multi-speaker baseline model (on an AWS g5.2xlarge instance) takes 5-7 days (100 epochs with 32 batch size and all 3 datasets (i.e., VCTK, LibriTTS_tc360, andpotion_Salut)).
|
||||
|
||||
1. At the end of a training run, there will be the following files in the result folder:
|
||||
|
||||
```txt
|
||||
results/baseline-models/vits_vctk-March-23-2022_03+43AM-0000000/
|
||||
|-- best_model.pth .................................... best model using avg_loss_0 (NOT the best model; suggest to ignore for now)
|
||||
|-- best_model_19096.pth .............................. same as best_model.pth (suggest to ignore for now)
|
||||
|-- checkpoint_300000.pth ............................. fifth last checkpoint
|
||||
|-- checkpoint_310000.pth ............................. fourth last checkpoint
|
||||
|-- checkpoint_320000.pth ............................. third last checkpoint
|
||||
|-- checkpoint_330000.pth ............................. second last checkpoint
|
||||
|-- checkpoint_340000.pth ............................. last checkpoint
|
||||
|-- config.json ....................................... configuration file
|
||||
|-- events.out.tfevents.1648007016.ip-172-31-83-225 ... event log for entire training run including eval samples and charts (view via tensorboard)
|
||||
|-- speakers.pth ...................................... speaker embeddings
|
||||
|-- trainer_0_log.txt ................................. training log
|
||||
|-- train_multispeaker_baseline_model.py .............. copy of the training script
|
||||
```
|
||||
|
||||
Use the event log to determine which of the checkpoints corresponds to the best model.
|
||||
|
||||
### Clone a Voice based on the Mutli-speaker Baseline Model
|
||||
|
||||
1. To clone a new voice, you need at least 10 voice samples (ideally 30). Those voice recordings (and their corresponding transcription files) have to be arranged as follows (and compressed into a `.tgz`, `.tbz` or `.zip` archive):
|
||||
|
||||
```txt
|
||||
VOICE_DATASET_PATH/txt/1/1_001.txt
|
||||
VOICE_DATASET_PATH/txt/1/1_002.txt
|
||||
VOICE_DATASET_PATH/txt/1/1_003.txt
|
||||
...
|
||||
VOICE_DATASET_PATH/txt/1/1_029.txt
|
||||
VOICE_DATASET_PATH/txt/1/1_030.txt
|
||||
VOICE_DATASET_PATH/wav48/1/1_001.wav
|
||||
VOICE_DATASET_PATH/wav48/1/1_002.wav
|
||||
VOICE_DATASET_PATH/wav48/1/1_003.wav
|
||||
...
|
||||
VOICE_DATASET_PATH/wav48/1/1_028.wav
|
||||
VOICE_DATASET_PATH/wav48/1/1_029.wav
|
||||
VOICE_DATASET_PATH/wav48/1/1_030.wav
|
||||
```
|
||||
|
||||
1. Next, pre-process audio recordings to fit the format of audio samples (i.e., sampling rate) and pre-compute speaker embeddings:
|
||||
|
||||
```sh
|
||||
(potion-voice_venv) $ python prepare_datasets.py --dataset_preset potion_voice_cloning --dataset_archive_path ~/datasets/potion\ Recordings/potion-voice\ recordings/user123.tgz
|
||||
|
||||
Commencing preparation of dataset for multi-speaker baseline model training:
|
||||
|
||||
+ Dataset preset: potion_voice_cloning
|
||||
+ Dataset : /home/[REDACTED_HOMEDIR_USERNAME_3]/datasets/potion Recordings/potion-voice recordings/user123.tgz
|
||||
+ Output path : results/datasets
|
||||
+ Sampling rate : 22050
|
||||
|
||||
>>> Extracting archive ...
|
||||
>>> Resampling audio files to 16000Hz ...
|
||||
Resampling the audio files...
|
||||
Found 30 files...
|
||||
100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 30/30 [00:00<00:00, 39.40it/s]
|
||||
Done !
|
||||
>>> Extracting original archive again (overwritting previously resampled files)...
|
||||
>>> Resampling audio files to 22050Hz ...
|
||||
Resampling the audio files...
|
||||
Found 30 files...
|
||||
100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 30/30 [00:00<00:00, 37.61it/s]
|
||||
Done !
|
||||
|
||||
Completed preparing voice dataset for multi-speaker baseline model training; generated asset locations are as follows:
|
||||
--> results/datasets/sr22050/user123
|
||||
--> results/datasets/sr22050/user123/speakers.pth
|
||||
|
||||
Done; bye.
|
||||
```
|
||||
|
||||
1. Finally, trigger voice cloning:
|
||||
|
||||
```sh
|
||||
(potion-voice_venv) $ python3 clone_voice.py [-h] --baseline_model_path BASELINE_MODEL_PATH --speaker_dataset_path SPEAKER_DATASET_PATH --speaker_embeddings_path SPEAKER_EMBEDDINGS_PATH [--output_path OUTPUT_PATH] [--batch_size BATCH_SIZE] [--max_epochs MAX_EPOCHS] [--use_cpu] [--output_format {txt,json}]
|
||||
|
||||
Code to clone a voice from a given set of voice samples and a multi-speaker baseline model
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--baseline_model_path BASELINE_MODEL_PATH
|
||||
Path to multi-speaker baseline model (VITS model)
|
||||
--speaker_dataset_path SPEAKER_DATASET_PATH
|
||||
Path to voice cloning dataset
|
||||
--speaker_embeddings_path SPEAKER_EMBEDDINGS_PATH
|
||||
Path to speaker's embeddings file
|
||||
--output_path OUTPUT_PATH
|
||||
Path to store trained / generated assets
|
||||
--batch_size BATCH_SIZE
|
||||
Batch size for training run
|
||||
--max_epochs MAX_EPOCHS
|
||||
Maximum number of epochs for training run
|
||||
--use_cpu Signal that CPU should be used even if a CUDA-device is available
|
||||
--output_format {txt,json}
|
||||
Output format; available choices include 'txt' for human readible text and 'json' for JSON formatting
|
||||
```
|
||||
|
||||
Using the default settings and 30 audio samples, cloning a new voice (on an AWS g5.2xlarge instance) takes about one hour.
|
||||
|
||||
1. At the end of a voice cloning run, there will be the following files in the result folder:
|
||||
|
||||
```txt
|
||||
results/cloned-voices/vits_potion_clone-December-28-2022_10+52AM-1327031/
|
||||
|-- best_model_365097.pth .................. best model using avg_loss_0 (save to use)
|
||||
|-- best_model.pth ......................... same as best_model_365097.pth
|
||||
|-- checkpoint_365200.pth .................. last checkpoint
|
||||
|-- clone_voice.py ......................... copy of the clone_voice script used in this run
|
||||
|-- config.json ............................ configuration file
|
||||
|-- events.out.tfevents.1672195961.rigel ... event log for entire voice cloning run including eval samples and charts (view via tensorboard)
|
||||
|-- speakers.pth ........................... speaker's embeddings file
|
||||
|-- trainer_0_log.txt ...................... training log file
|
||||
```
|
||||
|
||||
Use the event log to confirm that the best model is indeed giving the best outputs.
|
||||
|
||||
### Monitoring Training Progress
|
||||
|
||||
Using tensorboard / tensorboardX, training progress (for both, multi-speaker baseline training and voice cloning) can be monitored and evaluation samples can be accessed.
|
||||
|
||||
1. Ensure AWS Security Group settings (inbound) are set appropriamust include:
|
||||
|
||||
```txt
|
||||
HTTPS TCP 443 0.0.0.0/0
|
||||
Custom_TCP TCP 6006 0.0.0.0/0
|
||||
```
|
||||
|
||||
+ Server-side, launch the tensorboard service:
|
||||
|
||||
```sh
|
||||
(potion-voice_venv) $ tensorboard --logdir=./results/baseline-models/vits_vctk-March-07-2022_09+47AM-0000000/ --host 0.0.0.0
|
||||
TensorFlow installation not found - running with reduced feature set.
|
||||
|
||||
NOTE: Using experimental fast data loading logic. To disable, pass
|
||||
"--load_fast=false" and report issues on GitHub. More details:
|
||||
https://github.com/tensorflow/tensorboard/issues/4784
|
||||
|
||||
TensorBoard 2.8.0 at http://0.0.0.0:6006/ (Press CTRL+C to quit)
|
||||
```
|
||||
|
||||
+ Locally, point your preferred Web browser to <http://PUBLIC_IPv4_DNS:6006/>
|
||||
|
||||
### Minimise a Cloned Voice
|
||||
|
||||
To minimise the size of a trained model, run the followng script which removes optimiser and discriminator components from the model -- those are only required for training but not for inference:
|
||||
|
||||
```sh
|
||||
(potion-voice_venv) $ python3 minimize_cloned_voice_model.py [-h] --voice_model_asset_path VOICE_MODEL_ASSET_PATH [--voice_model_name VOICE_MODEL_NAME] [--voice_model_config_name VOICE_MODEL_CONFIG_NAME] [--minimise_suffix MINIMISE_SUFFIX] [--overwrite_assets] [--output_format {txt,json}]
|
||||
|
||||
Code to minimise (i.e., remove optimiser & discriminator) a cloned voice model
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--voice_model_asset_path VOICE_MODEL_ASSET_PATH
|
||||
Path to directory storing cloned voice model and the corresponding configuration and speaker files
|
||||
--voice_model_name VOICE_MODEL_NAME
|
||||
Name of the (best) cloned voice model
|
||||
--voice_model_config_name VOICE_MODEL_CONFIG_NAME
|
||||
Name of the config file for the cloned voice model
|
||||
--minimise_suffix MINIMISE_SUFFIX
|
||||
Suffix to be used for minimised model and its assets (i.e., new config file)
|
||||
--overwrite_assets Signal whether existing model assets should be overwritten or not (default: do not overwrite)
|
||||
--output_format {txt,json}
|
||||
Output format; available choices include 'txt' for human readible text and 'json' for JSON formatting
|
||||
```
|
||||
|
||||
1. Command-line output sample for output format option "txt":
|
||||
|
||||
```sh
|
||||
$ python3 minimize_cloned_voice_model.py --voice_model_asset_path results/cloned-voices/vits_potion_clone-December-28-2022_10+52AM-1327031/
|
||||
Minimising given voice model:
|
||||
|
||||
+ Cloned voice model file path : results/cloned-voices/vits_potion_clone-December-28-2022_10+52AM-1327031/best_model.pth
|
||||
+ Cloned voice model config file : results/cloned-voices/vits_potion_clone-December-28-2022_10+52AM-1327031/config.json
|
||||
|
||||
> Using model: vits
|
||||
> Setting up Audio Processor...
|
||||
[...]
|
||||
Completed minimising cloned voice model. The resulting (modified) assets can be found at:
|
||||
--> Minimised voice model path : results/cloned-voices/vits_potion_clone-December-28-2022_10+52AM-1327031/best_model_light.pth
|
||||
--> Minimised voice model config path: results/cloned-voices/vits_potion_clone-December-28-2022_10+52AM-1327031/config_light.json
|
||||
|
||||
Done; bye.
|
||||
```
|
||||
|
||||
### Scoring a Cloned Voice
|
||||
|
||||
1. To score a cloned voice, run the following command:
|
||||
|
||||
```sh
|
||||
(potion-voice_venv) $ python3 score_cloned_voice.py [-h] --voice_dataset_path VOICE_DATASET_PATH --voice_model_path VOICE_MODEL_PATH --voice_model_config_path VOICE_MODEL_CONFIG_PATH --speaker_embeddings_path SPEAKER_EMBEDDINGS_PATH [--temp_path TEMP_PATH] [--keep_temp] [--use_cpu] [--output_format {txt,json}]
|
||||
|
||||
Compute quality score for a given voice model (cloned voice) wrt. a given set of voice recordings (original voice))
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--voice_dataset_path VOICE_DATASET_PATH
|
||||
Path to set of voice recordings (original voice)
|
||||
--voice_model_path VOICE_MODEL_PATH
|
||||
Path to cloned voice model
|
||||
--voice_model_config_path VOICE_MODEL_CONFIG_PATH
|
||||
Path to config file for the cloned voice model
|
||||
--speaker_embeddings_path SPEAKER_EMBEDDINGS_PATH
|
||||
Path to speaker's embeddings file (i.e., pre-computed embeddings typically stored with the speaker's dataset)
|
||||
--temp_path TEMP_PATH
|
||||
Path to store temporary speech assets
|
||||
--keep_temp Signal that temporary assets used for scoring should not be deleted once done
|
||||
--use_cpu Signal that CPU should be used even if a CUDA-device is available
|
||||
--output_format {txt,json}
|
||||
Output format; available choices include 'txt' for human readible text and 'json' for JSON formatting
|
||||
```
|
||||
|
||||
1. Command-line output sample for output format option "txt":
|
||||
|
||||
```sh
|
||||
$ python3 score_cloned_voice.py --voice_dataset_path results/datasets/sr22050/michael/wav48/1/ --voice_model_path results/cloned-voices/vits_potion_clone-December-28-2022_01+09AM-1327031/best_model.pth --voice_model_config_path results/cloned-voices/vits_potion_clone-December-28-2022_01+09AM-1327031/config.json --speaker_embeddings_path results/datasets/sr22050/michael/speakers.pth
|
||||
Computing similarity score for a given voice model (cloned voice) wrt. a given set of voice recordings (original voice):
|
||||
|
||||
+ Original voice recordings path: results/datasets/sr22050/michael/wav48/1/
|
||||
+ Cloned voice model file path : results/cloned-voices/vits_potion_clone-December-28-2022_01+09AM-1327031/best_model.pth
|
||||
+ Cloned voice model config file: results/cloned-voices/vits_potion_clone-December-28-2022_01+09AM-1327031/config.json
|
||||
+ Speaker embeddings file : results/datasets/sr22050/michael/speakers.pth
|
||||
|
||||
+ CUDA availability : True
|
||||
+ Compute device used : cuda
|
||||
|
||||
+ No. of speakers : 1
|
||||
+ Speaker's names : ['VCTK_old_1']
|
||||
+ No. of embeddings : 30
|
||||
|
||||
> Using model: vits
|
||||
> Setting up Audio Processor...
|
||||
|
||||
Loaded the voice encoder model on cuda in 0.01 seconds.
|
||||
|
||||
Completed computing similarity score for the two sets of recordings. The resulting similarity score is:
|
||||
--> 0.9127510190010071
|
||||
|
||||
Done; bye.
|
||||
```
|
||||
|
||||
1. Command-line output sample for output format option "json":
|
||||
|
||||
```sh
|
||||
(potion-voice_venv)$ python3 score_cloned_voice.py --voice_dataset_path results/datasets/sr22050/michael/wav48/1/ --voice_model_path results/cloned-voices/vits_potion_clone-December-28-2022_01+09AM-1327031/best_model.pth --voice_model_config_path results/cloned-voices/vits_potion_clone-December-28-2022_01+09AM-1327031/config.json --speaker_embeddings_path results/datasets/sr22050/michael/speakers.pth --output_format json
|
||||
> Using model: vits
|
||||
> Setting up Audio Processor...
|
||||
[...]
|
||||
Loaded the voice encoder model on cuda in 0.01 seconds.
|
||||
{"success": true, "in": {"voice_dataset_path": "results/datasets/sr22050/michael/wav48/1/", "voice_model_path": "results/cloned-voices/vits_potion_clone-December-28-2022_01+09AM-1327031/best_model.pth"}, "out": {"score": 0.91}}
|
||||
```
|
||||
|
||||
## Usage Examples for Speech Synthesizing
|
||||
|
||||
1. To generate speech for a given cloned voice, run the following command:
|
||||
|
||||
```sh
|
||||
(potion-voice_venv) $ python3 synthesize_speech.py [-h] --voice_model_path VOICE_MODEL_PATH --voice_model_config_path VOICE_MODEL_CONFIG_PATH --speaker_embeddings_path SPEAKER_EMBEDDINGS_PATH --txt TXT [--output_path OUTPUT_PATH] [--target_sampling_rate TARGET_SAMPLING_RATE] [--speech_sample_wav_path SPEECH_SAMPLE_WAV_PATH] [--speech_sample_txt SPEECH_SAMPLE_TXT] [--trim_silence] [--use_cpu] [--output_format {txt,json}]
|
||||
|
||||
Code to synthesize speech for a given voice model
|
||||
|
||||
options:
|
||||
-h, --help show this help message and exit
|
||||
--voice_model_path VOICE_MODEL_PATH
|
||||
Path to cloned voice model
|
||||
--voice_model_config_path VOICE_MODEL_CONFIG_PATH
|
||||
Path to config file for the cloned voice model
|
||||
--speaker_embeddings_path SPEAKER_EMBEDDINGS_PATH
|
||||
Path to speaker's embeddings file (i.e., pre-computed embeddings typically stored with the speaker's dataset)
|
||||
--txt TXT Text to synthesize
|
||||
--output_path OUTPUT_PATH
|
||||
Path to store generated speech assets
|
||||
--target_sampling_rate TARGET_SAMPLING_RATE
|
||||
Desired sampling rate (in Hz) for output file
|
||||
--speech_sample_wav_path SPEECH_SAMPLE_WAV_PATH
|
||||
Path to a sample utterance of the speaker (used for style transfer)
|
||||
--speech_sample_txt SPEECH_SAMPLE_TXT
|
||||
Text of the sample utterance of the speaker (used for style transfer)
|
||||
--trim_silence Signal whether to trim silence from synthesised speech
|
||||
--use_cpu Signal that CPU should be used even if a CUDA-device is available
|
||||
--output_format {txt,json}
|
||||
Output format; available choices include 'txt' for human readible text and 'json' for JSON formatting
|
||||
```
|
||||
|
||||
1. Command-line output sample for output format option "txt":
|
||||
|
||||
```sh
|
||||
(potion-voice_venv) $ python3 synthesize_speech.py --voice_model_path results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/best_model.pth --voice_model_config_path results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/config.json --speaker_embeddings_path results/datasets/sr22050/[REDACTED_HOMEDIR_USERNAME_2]/speakers.pth --txt "Hi person_82, it works!"
|
||||
Commencing speech synthesizing:
|
||||
|
||||
+ Voice model file path : results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/best_model.pth
|
||||
+ Voice model config file: results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/config.json
|
||||
+ Speaker embeddings file: results/datasets/sr22050/[REDACTED_HOMEDIR_USERNAME_2]/speakers.pth
|
||||
+ Output path : results/speech
|
||||
+ Text to synthesize : Hi person_82, it works!
|
||||
|
||||
+ CUDA availability : True
|
||||
+ Compute device used : cuda
|
||||
+ No. of speakers : 1
|
||||
+ Speaker's names : ['VCTK_old_1']
|
||||
+ No. of embeddings : 30
|
||||
|
||||
> Using model: vits
|
||||
> Setting up Audio Processor...
|
||||
[...]
|
||||
>>> Saving original output to : results/speech/b4189e9e-6142-4dad-8577-6de77087ffd1.wav
|
||||
>>> Saving resampled output to: results/speech/b4189e9e-6142-4dad-8577-6de77087ffd1_sr48000.wav
|
||||
|
||||
Speech synthesizing has completed. Bye.
|
||||
```
|
||||
|
||||
1. Command-line output sample for output format option "txt":
|
||||
|
||||
```sh
|
||||
(potion-voice_venv) $ python3 synthesize_speech.py --voice_model_path results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/best_model_light.pth --voice_model_config_path results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/config_light.json --speaker_embeddings_path results/datasets/sr22050/[REDACTED_HOMEDIR_USERNAME_2]/speakers.pth --txt "Hi person_82, it works!" --output_format json
|
||||
> Using model: vits
|
||||
> Setting up Audio Processor...
|
||||
[...]
|
||||
{"success": true, "in": {"voice_model_path": "results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/best_model_light.pth", "voice_model_config_path": "results/cloned-voices/vits_potion_clone-December-28-2022_08+30AM-1327031/config_light.json", "speaker_embeddings_path": "results/datasets/sr22050/[REDACTED_HOMEDIR_USERNAME_2]/speakers.pth"}, "out": {"speech_original_path": "results/speech/c99e494c-e1f9-4c12-9095-255cf7db792b.wav", "speech_resampled_path": "results/speech/c99e494c-e1f9-4c12-9095-255cf7db792b_sr48000.wav"}}
|
||||
```
|
||||
|
||||
### Scoring a Synthesised Salutation
|
||||
|
||||
1. To score a synthesised salutation, run the following command:
|
||||
|
||||
```sh
|
||||
(potion-voice_venv)$ python3 score_salutation.py [-h] --recording_path RECORDING_PATH --first_name FIRST_NAME [--output_format {txt,json}]
|
||||
|
||||
Score a given salutation recording wrt. its desired content, the actual salutation recording, and a generated transcription (using Potion's internal Transciption API) of the recording.
|
||||
|
||||
optional arguments:
|
||||
-h, --help show this help message and exit
|
||||
--recording_path RECORDING_PATH
|
||||
Path to salutation recoding (.wav audio file)
|
||||
--first_name FIRST_NAME
|
||||
First name that the salutation recoding is meant to use
|
||||
--output_format {txt,json}
|
||||
Output format; available choices include 'txt' for human readible text and 'json' for JSON formatting
|
||||
```
|
||||
|
||||
1. Command-line output sample for output format option "txt":
|
||||
|
||||
```sh
|
||||
(potion-voice_venv)$ python3 score_salutation.py --recording_path /home/[REDACTED_HOMEDIR_USERNAME_3]/person_82_-_Hey_person_83.wav --first_name person_83
|
||||
Commencing scoring of the given salutation recording:
|
||||
|
||||
+ Salutation recording path: /home/[REDACTED_HOMEDIR_USERNAME_3]/person_82_-_Hey_person_83.wav
|
||||
+ Salutation first name : person_83
|
||||
|
||||
>> Salutation score : 0.892155
|
||||
|
||||
Done; bye.
|
||||
```
|
||||
|
||||
1. Command-line output sample for output format option "json":
|
||||
|
||||
```sh
|
||||
(potion-voice_venv)$ python3 score_salutation.py --recording_path /home/[REDACTED_HOMEDIR_USERNAME_3]/person_82_-_Hey_person_83.wav --first_name person_83 --output_format json
|
||||
{"in": {"recording_path": "/home/[REDACTED_HOMEDIR_USERNAME_3]/person_82_-_Hey_person_83.wav", "first_name": "person_83"}, "out": {"score": 0.89}}
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
1. How to better monitor GPU load / utilisation?
|
||||
|
||||
+ Install an interactive NVIDIA-GPU process viewer such as `nvitop`:
|
||||
|
||||
```sh
|
||||
$ python3 -m pip install nvitop
|
||||
Collecting nvitop
|
||||
[...]
|
||||
Installing collected packages: nvidia-ml-py, termcolor, psutil, nvitop
|
||||
Successfully installed nvidia-ml-py-11.495.46 nvitop-0.8.0 psutil-5.9.2 termcolor-2.0.1
|
||||
````
|
||||
|
||||
+ Run via command-line: `nvitop`:
|
||||
|
||||
```sh
|
||||
Tue Sep 13 02:09:07 2022
|
||||
╒═════════════════════════════════════════════════════════════════════════════╕
|
||||
│ NVITOP 0.8.0 Driver Version: 515.65.01 CUDA Driver Version: 11.7 │
|
||||
├───────────────────────────────┬──────────────────────┬──────────────────────┤
|
||||
│ GPU Name Persistence-M│ Bus-Id Disp.A │ Volatile Uncorr. ECC │
|
||||
│ Fan Temp Perf Pwr:Usage/Cap│ Memory-Usage │ GPU-Util Compute M. │
|
||||
╞═══════════════════════════════╪══════════════════════╪══════════════════════╪══════════════════════════╕
|
||||
│ 0 A10G On │ 00000000:00:1E.0 Off │ 0 │ MEM: █████████▊ 65.1% │
|
||||
│ 0% 47C P0 192W / 300W │ 14982MiB / 22.49GiB │ 100% Default │ UTL: ███████████████ MAX │
|
||||
╘═══════════════════════════════╧══════════════════════╧══════════════════════╧══════════════════════════╛
|
||||
[ CPU: ██████████▏ 18.1% ] ( Load Average: 1.07 1.11 1.04 )
|
||||
[ MEM: ███████████▎ 20.2% ] [ SWP: ▏ 0.3% ]
|
||||
|
||||
╒════════════════════════════════════════════════════════════════════════════════════════════════════════╕
|
||||
│ Processes: ubuntu@ip-172-31-95-84 │
|
||||
│ GPU PID USER GPU-MEM %SM %CPU %MEM TIME COMMAND │
|
||||
╞════════════════════════════════════════════════════════════════════════════════════════════════════════╡
|
||||
│ 0 2100 C ubuntu 14463MiB 90 103.7 9.6 5.4 days python3 train_multispeaker_baseline_model.py │
|
||||
╘════════════════════════════════════════════════════════════════════════════════════════════════════════╛
|
||||
```
|
||||
Reference in New Issue
Block a user