54 lines
1.4 KiB
Markdown
54 lines
1.4 KiB
Markdown
# TTS — Audio Generation for ASFR Stories
|
|||
|
|
|
||
|
|
## What
|
||
|
|
|
||
|
|
Generates audiobooks from story files using **Silero TTS v5** (Russian). Two voices:
|
||
|
|
- **Narrator** — male voice (`eugene`) for the story prose
|
||
|
|
- **Cyborg voice** — same voice with robotic effect (pitch shift + ring modulation + bitcrush) for system error messages (`ОШИБКА:`, `ПРЕДУПРЕЖДЕНИЕ:`)
|
||
|
|
|
||
|
|
## Setup
|
||
|
|
|
||
|
|
```bash
|
||
|
|
uv add torch torchaudio silero silero-stress soundfile librosa
|
||
|
|
```
|
||
|
|
|
||
|
|
## Usage
|
||
|
|
|
||
|
|
```bash
|
||
|
|
# Generate Silas audiobook
|
||
|
|
uv run tts/tts_silas.py Silas.md
|
||
|
|
|
||
|
|
# Generate any story
|
||
|
|
uv run tts/tts_silas.py Logan.md
|
||
|
|
```
|
||
|
|
|
||
|
|
Output goes to `tts/output/`:
|
||
|
|
- `<name>_full.wav` — complete audiobook
|
||
|
|
- `<name>_demo.wav` — first 4 segments (~1 min)
|
||
|
|
|
||
|
|
## Voice Options
|
||
|
|
|
||
|
|
Available Silero v5_ru voices:
|
||
|
|
- `aidar` — deep male (can be monotone)
|
||
|
|
- `eugene` — male, more expressive intonation ← default
|
||
|
|
- `baya` — female, soft
|
||
|
|
- `kseniya` — female, clear
|
||
|
|
- `xenia` — female, calm
|
||
|
|
|
||
|
|
Change `SPEAKER` in `tts_silas.py` to switch voices.
|
||
|
|
|
||
|
|
## Robotic Effect
|
||
|
|
|
||
|
|
System messages get these audio effects:
|
||
|
|
- Pitch shift: -4 semitones (deeper, more mechanical)
|
||
|
|
- Ring modulation: 50 Hz carrier (metallic timbre)
|
||
|
|
- Bitcrush: 200 levels (digital distortion)
|
||
|
|
|
||
|
|
Parameters in `make_robotic()` can be tuned.
|
||
|
|
|
||
|
|
## Known Issues
|
||
|
|
|
||
|
|
- Silero can misplace Russian word stress on complex sentences
|
||
|
|
- Long paragraphs are split at sentence boundaries
|
||
|
|
- Robotic effect uses librosa which has deprecation warnings (harmless)
|