Speech recognition with XPENG-AI/X-AuT (Qwen3-ASR-0.6B with a 14-layer pruned audio encoder), running fully in your browser on WebGPU. Audio never leaves the device.
Weights download once, then load from cache.
or drop a file here
Examples
Transcript appears here.
About this demo
ONNX export: mrfakename/X-AuT-ONNX. The 14-layer audio encoder (8-bit weights) and the 28-layer Qwen3 decoder (int4 k-quant with int8 for sensitive layers) run on WebGPU via onnxruntime-web; the KV cache stays on the GPU. fp16 matches PyTorch output; q4f16 is within ~0.4 WER points of it on our test set.
Long recordings are split into ≤30 s pieces at quiet points. Auto-detect lets the model emit its own language tag; picking a language forces plain transcription in that language.
The context box is passed as the system prompt, which Qwen3-ASR uses for contextual biasing (names, terms). X-AuT was finetuned mainly on Chinese and English.
Weights are CC-BY-NC-4.0 (non-commercial). Example clips: LibriSpeech and FLEURS (CC-BY-4.0).