To ensure optimal model performance, your audio must meet these specifications:
โ
Sample Rate: 16 kHz or higher (e.g., 16kHz, 44.1kHz, 48kHz)
โ
Bit Depth: 16-bit or higher quality (16-bit, 24-bit, 32-bit float)
Supported Formats: WAV, FLAC, MP3, M4A
Processing: Higher sample rates automatically downsampled to 16kHz | Higher bit depths quantized to 16-bit to match training data
Audio below these specifications will be rejected to maintain model accuracy.