Tossed distorted audio samples to an open-weight voice model; it did fairly well.
A user stress-tested open-weight voice model Confucius4 with World Cup commentary clips—screaming, held-breath explosions, and a goalkeeper's shaky post-match interview. The model translates directly from audio, not from a transcript, and preserved short emotional bursts well. Long sentences degraded into synthetic quality because the model has to guess the rest mid-sentence. The post doesn't disclose model size, training data, or Chinese support.