Stream audio from the mic to the model to the speaker
  • Python 74.7%
  • Rust 25.3%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Brandon Thomas 4a222ad815 realtime-voice-conversion based on last year's work
This is based on the work I did in (July - October)ish 2019 to get real
time voice conversion working. I've removed it from the `voder` library
and made it into an isolable program. I'll work on extending the
capabilities.
2020-09-11 09:06:48 -04:00
cycle_gan realtime-voice-conversion based on last year's work 2020-09-11 09:06:48 -04:00
protos realtime-voice-conversion based on last year's work 2020-09-11 09:06:48 -04:00
src realtime-voice-conversion based on last year's work 2020-09-11 09:06:48 -04:00
.gitignore realtime-voice-conversion based on last year's work 2020-09-11 09:06:48 -04:00
build.rs realtime-voice-conversion based on last year's work 2020-09-11 09:06:48 -04:00
Cargo.lock realtime-voice-conversion based on last year's work 2020-09-11 09:06:48 -04:00
Cargo.toml realtime-voice-conversion based on last year's work 2020-09-11 09:06:48 -04:00
README.md realtime-voice-conversion based on last year's work 2020-09-11 09:06:48 -04:00

Voice Converter Sidecar

This project spins up a Rust binary to handle audio input and output, then ships microphone audio over proto/zeromq to a Python sidecar running CycleGAN-VC to convert it into target speech.

CycleGAN

I've included a pared down copy of CycleGAN for the purposes of model evaluation. It handles all of the audio buffering / sidecar integration pieces. The original CycleGAN I used is here.

Note on Training

It's possible to train and evaluate at the same time using dual GPUs (at least my dual 1080Ti setup). To run the sidecar on a particular GPU (0-indexded), use:

CUDA_VISIBLE_DEVICES=1 ./sidecar.py

Proto Compilation (for Python)

Codegen for Rust is built in. Codegen for Python uses,

protoc --python_out=cycle_gan protos/audio.proto

Current Results

Currently there is 4.39 seconds of delay between speaking and generated output with the sidecar setup on my desktop computer. This is really great and seems promising.

This gets up to 6.0 seconds later. Drift continues to accrue, but it's a slow build.