TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models
Abstract
Unified audio models that can both understand and generate audio are proliferating, yet a basic question goes unasked: do a model's two heads agree about the same audio? Current practice grades understanding, generation, and editing in isolation on separate benchmarks, and never checks whether a model can make sense of its own generations. We present TORUS, the first self-coherence test for audio-native unified models. TORUS comprises 48 three-stage self-coherence tests carrying 432 six-option questions spanning speech, sound, and music across five task families. Each test runs a model through generation, an edit, and a counterfactual edit, and after every stage the same model answers spec-keyed questions about its own clip - with only the audio in context. We evaluate five open unified models alongside a Cascaded Baseline of state-of-the-art specialist generation, editing, and understanding models. The best unified model answers 50.5 percent of questions against the baseline's 63.2 percent and a 16.7 percent chance floor. Models degrade once editing begins and struggle most on physical causality. Our results reveal a general lack of self-coherence across audio models, unified and specialist alike, and position self-coherence as an essential test for future audio systems.
Benchmark composition
TORUS comprises 48 three-stage tests, balanced across speech, sound, and music at 16 tests each, carrying 432 six-option questions across five task families.
48 tests by audio type
48
tests
Speech16 tests 33%
Sound16 tests 33%
Music16 tests 33%
432 questions by task family
432
questions
Source composition63 q 15%
Physical causality54 q 12%
Scene acoustics90 q 21%
Temporal ordering135 q 31%
Interaction & congruence90 q 21%
Five task families
ASource composition.Which sources are present and what they are - the identities the clip must contain.
BPhysical causality.The coupling between a cause and its acoustic consequence - what an action does to the scene.
CScene acoustics.How the scene is rendered - distance, room, timbre, and spatial placement.
DTemporal ordering.Order, timing, counts, and repetition across the clip over time.
EInteraction dynamics and semantic congruence.How sources relate and respond, and whether meaning and affect line up.
Worked examples
This is what a PDF cannot do. The paper prints one test as a figure; here you can listen to all six systems generate and edit audio, then see how each one answered questions about its own output. Pick a tile to jump to a full test.
In each question the key is highlighted; the five other options are plausible-sounding wrong answers. The per-test matrices show texture (what failure sounds and looks like); the aggregate claims live in the tables below. Chance is 1 in 6.
soundB physical causalityThe same shove taken by the wall
What to noticeAnatomy of a test. Every model holds the opening scene, then slips once the edit changes the physics. Watch Stage 3, where the weaker models answer as if the earlier drawer-slam were still the scene.
Stage 1 - GenerateT2A
promptIn a kitchen with a running refrigerator, a person gently pushes a wooden cutlery drawer shut.
intended audioA kitchen, faint fridge hum. A wooden cutlery drawer is pushed shut gently: a short soft roll on its runners, a light click of the catch, one faint chink.
Stage 2 - EditStyle Transfer
promptShove the drawer shut hard instead of pushing it gently.
intended audioThe same drawer is shoved shut hard: a fast rumbling roll, a loud wooden bang at the stop, the catch snapping, and the cutlery jangling for a second before settling.
Stage 3 - Counterfactual editStyle Transfer
promptThe cabinet stands on castors; push the entire cabinet back so it rolls and hits the wall, closing the drawer.
intended audioThe cabinet stands free on castors: it trundles back and hits the wall with a heavy thud and a low booming ring, the drawer closing with a soft knock.
Show all nine questions (three per stage)
Stage 1 - Generate
S1.1What is the primary action and setting depicted in this recording?
A wooden sliding compartment is gently pushed shut in a room with a faint background hum.
A metal filing cabinet drawer is slammed shut in a noisy office environment.
A wooden sliding compartment is pulled fully open in a completely silent room.
A heavy wooden door is slowly latched shut in an outdoor setting with wind noise.
The sliding compartment is 30 centimeters wide and painted white, operating in a modern kitchen.
A plastic storage bin is dropped onto a carpeted floor in a quiet room.
S1.2In what sequence do the sounds associated with the sliding compartment occur?
A brief, soft rolling sound occurs first, followed immediately by a quiet click and a single faint metallic chink.
A loud metallic crash occurs first, followed by a long, slow rolling sound and a final heavy thud.
A sharp metallic chink occurs first, followed by a soft rolling sound and a final quiet click.
A quiet click occurs first, followed by a metallic chink, with the soft rolling sound occurring last.
The user pushes the compartment, then watches it slide, and finally checks if it is locked.
A continuous metallic rattling occurs throughout, with no distinct rolling or clicking sounds.
S1.3What is the manner of the closing event for the sliding compartment?
The movement is slow and low-energy, resulting in a very soft, cushioned contact sound.
The movement is extremely rapid and forceful, resulting in a loud, sharp wooden impact.
The movement is jerky and hesitant, resulting in multiple squeaks and uneven scraping sounds.
The movement is completely silent, with only the final metallic contact being audible.
The sliding compartment is pushed by a person who is in a hurry, using their elbow to close it.
The movement is moderate in speed, resulting in a hollow, echoing wooden bounce.
Stage 2 - Edit
S2.1What primary event occurs in this recording?
A sliding compartment is violently slammed shut, producing a loud wooden impact and a sharp metallic rattle.
A sliding compartment is gently pushed shut, producing a soft click and a faint metallic chink.
A heavy wooden door is slammed shut, causing a nearby window pane to rattle loudly.
A sliding compartment is pulled open quickly, causing its contents to slide to the back with a dull thud.
A person angrily shoves a wooden drawer using maximum physical force to secure the contents.
A wooden box is dropped from a height, causing it to splinter and scatter loose metal pieces.
S2.2What is the sequence of acoustic events during the closing of the sliding compartment?
A rapid, heavy rumbling sound is followed by a loud wooden bang, which then triggers a brief period of jangling metal.
A loud wooden bang occurs first, followed by a rapid rumbling sound and a final quiet metallic click.
A brief, soft rolling sound occurs first, followed immediately by a quiet click and a single faint metallic chink.
A sharp metallic jangling occurs first, followed by a heavy rumbling sound and a final soft wooden knock.
The hand contacts the handle, the drawer slides along its metal tracks, and the internal cutlery is displaced.
A continuous wooden scraping sound occurs first, ending abruptly with no metallic sounds at all.
S2.3What is the manner of the sound made by the loose items inside the sliding compartment during the event?
The loose items jangle loudly and chaotically for about a second before settling.
The loose items produce only a single, barely audible metallic chink.
The loose items remain completely silent and undisturbed throughout the event.
The loose items slide smoothly with a continuous, low-pitched scraping sound.
The loose items consist of stainless steel forks and spoons that are slightly scratched from use.
The loose items rattle in a rhythmic, repeating pattern that slowly fades out over several seconds.
Stage 3 - Counterfactual edit
S3.1Which of these describes the events present in this recording?
A heavy rolling structure trundles and collides with a surface, accompanied by a separate, quieter sliding compartment closing sound.
A heavy rolling structure trundles and collides with a surface, but no internal sliding compartment movement or closing sound occurs.
A heavy rolling structure is pushed smoothly across a carpeted floor without colliding or causing any other sounds.
A sliding compartment is slammed shut violently, with no rolling sound from the larger structure.
A wooden cabinet on four rubber castors is pushed two meters across a kitchen floor into a plaster wall.
A heavy object is dropped directly onto a wooden floor, followed by the sound of a sliding compartment opening.
S3.2What is the sequence of acoustic events in this recording?
A low trundling sound leads to a heavy, resonant thud, followed shortly after by a soft, dull knock.
A soft, dull knock occurs first, followed by a low trundling sound and a final heavy, resonant thud.
A rapid, heavy rumbling sound is followed by a loud wooden bang, which then triggers a brief period of jangling metal.
A heavy, resonant thud occurs first, followed immediately by a continuous, high-pitched metallic rattling.
The operator pushes the cabinet, the wheels rotate on the floor, and the drawer slides shut due to inertia.
A low trundling sound occurs continuously, with a soft knock and a heavy thud happening at the exact same instant.
S3.3What is the manner of the sound produced by the sliding compartment closing during this event?
The sliding compartment closes with a quiet, soft knock and no metallic jangling.
The sliding compartment closes with a loud, sustained metallic jangling of loose items.
The sliding compartment closes with a sharp, high-pitched wooden snap and a loud click.
The sliding compartment closes with a long, screeching scrape followed by a heavy crash.
The sliding compartment is made of oak and contains heavy silverware that remains perfectly organized.
The sliding compartment closes with a muffled, rubbery thud that repeats twice.
Model
Stage 1
Stage 2
Stage 3
Cascaded Baseline
Audex-30B
Audio-Omni
Audex-2B
UniAudio-2
Unified-IO 2
Model
S1.1
S1.2
S1.3
S2.1
S2.2
S2.3
S3.1
S3.2
S3.3
Score
Cascaded Baseline
✓
✓
✓
✗
✗
✓
✓
✗
✗
5/9
Audex-30B
✓
✓
✗
✓
✗
✓
✓
✗
✗
5/9
Audio-Omni
✓
✗
✗
✓
✗
✓
✓
✗
✗
4/9
Audex-2B
✗
✓
✗
✗
✗
✓
✗
✗
✗
2/9
UniAudio-2
✗
✓
✗
✗
✗
✓
✗
✗
✗
2/9
Unified-IO 2
✗
✓
✗
✓
✗
✓
✗
✗
✗
3/9
speechD temporal structureStage line paced out, then run on one breath
What to noticeSpeech narrows the gap. On the temporal order of spoken events, the best unified model (Audio-Omni, 6 of 9) trails the baseline (8 of 9) by far less than on sound or music.
Stage 1 - GenerateT2A
promptIn a rumbling hall, an actor projects: 'The messenger arrived at dawn and nobody opened the gate.', inserting three long silences within the line.
intended audioA hall with a long tail and low air rumble. An actor projects 'The messenger arrived at dawn and nobody opened the gate.' with three long silences inside it.
Stage 2 - EditStyle Transfer
promptMake the actor deliver the line three times faster without any pauses.
intended audioA hall with a long tail and low air rumble. The actor runs the same line straight through on one breath, three times faster, the tails washing over the words.
Stage 3 - Counterfactual editStyle Transfer
promptModify the actor's voice to become weak, breathy, and strained during the final third of the line.
intended audioA hall with a long tail and low air rumble. Two thirds through the fast unbroken line, the voice thins to a weak, breathy, strained sound for the last words.
Show all nine questions (three per stage)
Stage 1 - Generate
S1.1What is the primary sound source and the environment heard in this recording?
A single voice speaking a sentence inside a highly reverberant, rumbling hall.
A single voice speaking a sentence in a completely dry, soundproofed room with no echo.
A group of voices chanting in unison inside a small, tiled bathroom.
A text description of a person speaking in a large, empty warehouse.
A single voice whispering a sentence outdoors with wind blowing in the background.
A synthesized robotic voice repeating a sentence inside a small carpeted office.
S1.2How do the spoken words and the background environment alternate over time?
The spoken phrases are separated by long pauses where only the room's reverberation and rumble continue.
The spoken phrases are delivered continuously without any pauses, followed by a long silence at the very end.
The background rumble plays first for several seconds, followed by the entire sentence spoken without any gaps.
The text transcript shows punctuation marks indicating where the speaker should have paused.
The spoken phrases overlap with each other, creating a continuous layer of speech with no silence at all.
The speech starts with a long pause, and then each word is followed by an immediate loud echo that cuts off the next word.
S1.3Which of the following describes the pacing and structure of the spoken sentence?
The sentence is delivered slowly, interrupted by three distinct, long pauses.
The sentence is delivered at a normal conversational pace with no pauses at all.
The sentence is delivered extremely fast, with brief gaps after every single word.
The sentence contains exactly twelve words spoken in a highly dramatic tone.
The sentence is delivered slowly, with a single long pause right in the middle of the line.
The sentence is delivered with a constantly accelerating pace, ending in a rapid rush of words.
Stage 2 - Edit
S2.1What happens in this scene?
The speaker delivers the entire sentence rapidly in a single continuous breath without any pauses.
The speaker delivers the sentence slowly, with three long pauses separating the phrases.
The speaker delivers the sentence at a normal, steady pace with brief pauses between words.
The transcript of the sentence is displayed on a screen with a fast-forward icon next to it.
The speaker whispers the sentence slowly, pausing after every word to take a deep breath.
The speaker delivers only the first half of the sentence and then stops completely.
S2.2How does the sound of the speech interact with the room's acoustics over time?
The speech runs continuously, causing the reverberant tail of earlier words to overlap and blend with the following words.
Each word is followed by a distinct echo that completely dies down before the next word begins.
The speech is interrupted by long silences that allow the reverberant tail of each phrase to fully decay to silence.
The room acoustics are described in the audio metadata as having a three-second decay time.
The reverberation gets quieter and quieter as the speaker talks faster, ending in a dry sound.
The speech starts dry and the reverberation only appears as a single loud burst after the final word.
S2.3What is the acoustic character of the words as a result of the delivery style in this space?
The words are partially washed over and blended together by the accumulating reverberation.
The words remain perfectly crisp, dry, and distinct from one another.
The words are separated by clear silence, keeping each phrase completely distinct and sharp.
The words are written in a cursive font where the letters connect continuously.
The words sound metallic and robotic due to digital time-stretching artifacts.
The words are completely unintelligible because of a loud, high-pitched feedback loop.
Stage 3 - Counterfactual edit
S3.1Which of these is present in this scene?
A rapid, continuous delivery of the spoken sentence inside a highly reverberant hall.
A slow delivery of the spoken sentence with three long, silent pauses.
A normal-paced delivery of the sentence in a completely dry room with no echo.
A text prompt asking the speaker to read the sentence quickly in a large room.
A rapid delivery of the sentence accompanied by a loud, rhythmic drum beat.
A slow, whispered reading of a completely different sentence in an outdoor environment.
S3.2How does the quality of the voice change over the course of the spoken sentence?
The voice projects strongly for the first two-thirds of the sentence, then shifts to a weak, strained quality for the final words.
The voice maintains a perfectly consistent, strong projection from the very first word to the last.
The voice starts weak and strained, but gradually becomes stronger and more projected by the end of the sentence.
The volume meter on the screen shows a sudden drop in decibels at the exact midpoint of the clip.
The voice alternates between strong projection and a weak whisper with every alternating word.
The voice starts strong, turns into a whisper in the middle, and ends with a loud scream.
S3.3What is the vocal character of the final words heard at the end of the sentence?
The final words are spoken with a thin, breathy, and strained vocal quality.
The final words maintain the same strong, projected, and resonant vocal quality as the beginning.
The final words are spoken in a deep, booming, and highly authoritative bass voice.
The final words are written in a smaller, italicized font to suggest a quieter tone.
The final words are replaced by a completely different speaker with a high-pitched voice.
The final words are spoken with a heavy electronic vocoder effect.
Model
Stage 1
Stage 2
Stage 3
Cascaded Baseline
Audex-30B
Audio-Omni
Audex-2B
UniAudio-2
Unified-IO 2
Model
S1.1
S1.2
S1.3
S2.1
S2.2
S2.3
S3.1
S3.2
S3.3
Score
Cascaded Baseline
✓
✓
✗
✓
✓
✓
✓
✓
✓
8/9
Audex-30B
✗
✓
✓
✗
✓
✓
✓
✗
✗
5/9
Audio-Omni
✓
✓
✓
✗
✓
✗
✓
✗
✓
6/9
Audex-2B
✗
✗
✓
✗
✗
✓
✓
✗
✓
4/9
UniAudio-2
✓
✓
✓
✓
✗
✗
✓
✗
✗
5/9
Unified-IO 2
✗
✗
✗
✗
✗
✗
✗
✗
✓
1/9
musicA scene contentBanjo joined by a harmonium drone that goes flat
What to noticeMusic is where unified models collapse. The specialist baseline answers 9 of 9, yet UniAudio-2 and Unified-IO 2 score 0 of 9 on clips they generated themselves.
Stage 1 - GenerateT2M
promptA clawhammer banjo plays a moderate rolling pattern close-mic'd in a small room, with rain pattering on the window.
intended audioIn a small room with rain pattering on the window, a clawhammer banjo plays a rolling pattern at a moderate tempo, close to the microphone. Nothing else plays.
Stage 2 - EditAdd
promptAdd a harmonium playing a low sustained note underneath the stringed instrument.
intended audioA harmonium enters underneath on one low sustained note, in tune with the banjo, and holds it unbroken to the end. The two sit together with no beating.
Stage 3 - Counterfactual editStyle Transfer
promptSlightly detune the drone instrument against the stringed instrument.
intended audioThe harmonium note is now tuned slightly flat against the banjo, so their combined tone swells and thins about three times a second in a slow regular pulsing.
Show all nine questions (three per stage)
Stage 1 - Generate
S1.1Which combination of sounds is heard in this recording?
A stringed instrument playing a pattern while weather sounds patter in the background.
A stringed instrument playing alone in a completely quiet, dry indoor room.
A wind instrument playing a melody while water splashes nearby.
A stringed instrument playing while a heavy thunderstorm with loud wind rages.
A person describing a rainy day while strumming an acoustic guitar.
A percussion instrument tapping a rhythm while a steady stream of water runs from a faucet.
S1.2What is the temporal relationship between the weather sounds and the stringed instrument?
The weather sounds patter continuously throughout the entire clip while the stringed instrument plays.
The weather sounds are heard only at the very beginning before the stringed instrument starts playing.
The stringed instrument plays a solo introduction, and the weather sounds only begin halfway through.
The two sounds alternate, with the instrument pausing whenever the weather sounds become louder.
A voice announces that it is raining, followed by a short pause, and then the instrument plays.
The weather sounds and the instrument play together for a brief moment and then both fade to silence early.
S1.3How is the primary stringed instrument played in this clip?
It is played at a moderate tempo using a rolling, repetitive pattern.
It is played with slow, widely spaced individual chords that ring out and decay.
It is played at an extremely fast, frantic speed with erratic, non-repetitive notes.
It is played using a slide technique, creating long, gliding pitches between notes.
A voice in the background counts the beats of the rhythm out loud.
It is played with a bow, producing sustained, scraping tones rather than plucked notes.
Stage 2 - Edit
S2.1What happens in this scene?
A low, sustained reed-like note joins underneath the stringed instrument and holds to the end.
Only the stringed instrument and the weather sounds are heard, with no background drone present.
A low, sustained bass guitar note plays a walking bassline underneath the stringed instrument.
A high-pitched, whistling wind sound enters and mimics the melody of the stringed instrument.
A voice says 'add drone' right before a low hum begins playing.
A sudden, loud percussion crash occurs, followed by complete silence from all instruments.
S2.2What is the temporal sequence of events in this clip?
The stringed instrument and weather sounds start first, and then a low sustained note enters and continues.
The low sustained note starts first as an introduction, and the stringed instrument enters later.
The stringed instrument and weather sounds play continuously from start to finish without any other sound entering.
The low sustained note plays only during the brief pauses of the stringed instrument, alternating with it.
The clip begins with a voice announcing the entry of the background note before it plays.
All sounds start together, but the low sustained note cuts out halfway through while the others continue.
S2.3How does the sustained background note sound in relation to the stringed instrument?
It blends smoothly and remains perfectly in tune, with no acoustic beating or pulsing.
It is completely out of tune, creating a harsh, clashing dissonance throughout the clip.
It rapidly slides up and down in pitch like a siren, never settling on a single note.
It stutter-steps, rapidly turning on and off like a gate effect rather than holding smoothly.
A voice in the background comments that the instruments are playing in harmony.
It sounds muffled and distorted, as if playing through a low-quality telephone line.
Stage 3 - Counterfactual edit
S3.1Which of these is present in this scene?
The stringed instrument and the background weather sounds are both present alongside a sustained low note.
A sustained low note plays alongside the weather sounds, but the stringed instrument is entirely missing.
The stringed instrument plays alongside a sustained low note, but the weather sounds have disappeared.
Only a solo sustained low note plays, with all other instruments and environmental sounds removed.
A voice describes a stringed instrument and rain, but only a low hum is actually heard.
An entirely new wind instrument plays alone, with no stringed instrument, weather, or low note present.
S3.2What is the temporal sequence of events in this clip?
The stringed instrument and weather sounds begin, and then a low note joins them and begins to pulse regularly.
The pulsing low note starts first, and then suddenly stops the moment the stringed instrument begins.
The low note pulses from the very first second, while the stringed instrument only enters at the very end.
The stringed instrument plays first, followed by a period of silence, and then the pulsing note plays alone.
A voice counts 'one, two, three' to mark the pulses before the stringed instrument starts.
All sounds start together, but the pulsing note immediately speeds up until it becomes a high squeal.
S3.3How does the combined sound of the instruments behave in this clip?
The combined tone swells and thins regularly, creating a slow, steady pulsing effect.
The background note remains completely smooth and steady, blending with the other instrument without any pulsing.
The pitch of the background note slides wildly and randomly up and down without any regular rhythm.
The volume of the entire clip suddenly drops to near-silence and then spikes to maximum volume once.
A voice in the background says 'pulsing' repeatedly in time with the music.
The instruments play with a harsh distortion that makes them sound fuzzy and cracked.
Model
Stage 1
Stage 2
Stage 3
Cascaded Baseline
Audex-30B
Audio-Omni
Audex-2B
UniAudio-2
Unified-IO 2
Model
S1.1
S1.2
S1.3
S2.1
S2.2
S2.3
S3.1
S3.2
S3.3
Score
Cascaded Baseline
✓
✓
✓
✓
✓
✓
✓
✓
✓
9/9
Audex-30B
✗
✓
✓
✗
✗
✗
✓
✗
✓
4/9
Audio-Omni
✗
✓
✗
✓
✗
✓
✓
✗
✗
4/9
Audex-2B
✗
✓
✗
✗
✗
✗
✓
✗
✗
2/9
UniAudio-2
✗
✗
✗
✗
✗
✗
✗
✗
✗
0/9
Unified-IO 2
✗
✗
✗
✗
✗
✗
✗
✗
✗
0/9
musicC rendering and spaceGuitar and metronome through an open then shut door
What to noticeScene acoustics in music. The baseline leads 7 of 9; every unified model lands between 0 and 4 - they render the notes but cannot read back the space.
Stage 1 - GenerateT2M
promptIn a carpeted room, a nylon-string guitar fingerpicks a slow arpeggio beside a ticking metronome, heard through an open doorway from another room with a humming fridge.
intended audioIn a carpeted room a nylon-string guitar fingerpicks a slow arpeggio while a metronome ticks. Heard from the next room through an open doorway. A fridge hums quietly.
Stage 2 - EditStyle Transfer
promptShut the door separating the listener from the instrument and the timekeeper.
intended audioThe connecting door is now shut. The guitar and metronome are quiet and dull, the string attack softened to a thud and each tick reduced to a soft knock.
Stage 3 - Counterfactual editStyle Transfer
promptMove the instrument and the timekeeper into a small tiled bathroom.
intended audioThe guitar and metronome are now in a small tiled bathroom. A reverberant tail follows each pluck and runs them together, and every knock is trailed by a splashy ring.
Show all nine questions (three per stage)
Stage 1 - Generate
S1.1What primary sound sources are active in this recording?
A plucked string instrument playing a slow melody, accompanied by a steady mechanical ticking and a low, continuous hum.
A piano playing a fast melody, accompanied by a digital beep and a loud, rushing wind sound.
A bowed string instrument playing a sustained note, accompanied by a clicking clock and a high-pitched whistle.
A voice describing a guitar playing next to a metronome and a refrigerator.
A brass instrument playing short notes, accompanied by a dripping faucet and a low rumbling engine.
A percussion mallet striking wooden bars, accompanied by a crackling fire and a steady fan hum.
S1.2How do the sounds in this recording overlap in time?
The low hum is continuous throughout, while the individual plucks and ticks occur repeatedly and concurrently.
The plucks and ticks play first in complete silence, followed by the low hum starting only after they stop.
The low hum and the plucks alternate back and forth, with the ticks occurring only during the silent gaps.
A voiceover announces each pluck and tick in order, while the low hum plays in the background.
The ticking occurs continuously at the start, then stops completely when the plucks and low hum begin.
All three sounds start and stop at the exact same moment, separated by long periods of total silence.
S1.3Which of the following describes the acoustic character of the plucked instrument and the ticking source?
They sound distant and indirect, as if heard from an adjacent room through an open passage.
They sound extremely close and direct, as if the listener is positioned right next to them in a dead space.
They sound highly reverberant and echoing, as if playing inside a large, empty stone cathedral.
A spoken narrator describes the distance and layout of the rooms.
They sound heavily distorted and fuzzy, as if played through a low-quality radio transmitter.
They sound bright and sparkling, with a sharp, crisp presence directly in the center of the stereo field.
Stage 2 - Edit
S2.1What is heard in this scene?
The plucks and ticks are extremely muffled and quiet, with their sharp high-frequency details completely lost.
The plucks and ticks are clearly audible and bright, sounding as if they are in the next room with an open doorway.
The plucks and ticks are completely silent, leaving only the low hum audible in the room.
The plucks and ticks are loud and distorted, sounding as if they are being played through a clipping speaker.
A voice states that the door has been shut and the sound is now muffled.
The plucks and ticks are accompanied by a loud, rushing wind sound that masks their high frequencies.
S2.2What is the temporal relationship between the sounds in this clip?
Each softened pluck occurs in sequence with a dull tick, both superimposed over a barely perceptible low-frequency hum.
The low-frequency hum plays alone for the first half, followed by a single loud pluck and tick at the very end.
The bright plucks and ticks alternate rapidly, while the low-frequency hum is completely absent.
The softened plucks and dull ticks occur only when the low-frequency hum momentarily cuts out.
A voice counts the number of plucks and ticks in order over a continuous hum.
The dull ticks occur in a rapid, continuous burst, followed by a slow sequence of softened plucks.
S2.3Which of the following describes the specific quality of the plucked instrument and the ticking source?
The plucks are reduced to soft, low-frequency thuds and the ticks sound like muted, hollow knocks.
The plucks sound sharp and metallic, while the ticks sound like high-pitched electronic beeps.
The plucks have a bright, crisp string attack and the ticks have a sharp, wooden click.
The plucks sound like wet splashes and the ticks sound like heavy dripping water.
A narrator explains that the high frequencies have been filtered out of the instruments.
The plucks sound like scratchy scrapes and the ticks sound like static crackles.
Stage 3 - Counterfactual edit
S3.1Which of these is present in this scene?
Both the plucked instrument and the ticking source are still active and playing together.
Only the plucked instrument is heard playing in the space, with the ticking source completely absent.
A completely different wind instrument is playing alone in a dry, carpeted room.
Only the ticking source is present, while the plucked instrument has been replaced by a continuous hum.
A voice lists the instruments that are currently active in the room.
Both the plucked instrument and the ticking source are completely silent, with only a loud rushing water sound present.
S3.2How do the sounds and their decays interact over time?
Each pluck and tick is immediately followed by a prolonged, overlapping decay that blurs into the next event.
Each pluck and tick decays instantly to silence before the next sound begins, keeping them completely separate.
The plucks and ticks are quiet and dry, with no decay or overlapping sound between them.
The decay tail plays first as a continuous wash, and the plucks and ticks occur only after the decay fades out.
A voice describes the length and overlap of the echo tails in seconds.
The decay of the plucks occurs only on the left channel, while the decay of the ticks occurs only on the right channel.
S3.3Which of the following describes the acoustic character of the plucked instrument and the ticking source?
The plucks and ticks are highly reverberant, each followed by a bright, metallic, and splashy ringing tail.
The plucks and ticks remain quiet, dull, and completely dry with no resonant decay.
The plucks and ticks sound completely dry and close-up, as if recorded in a heavily padded vocal booth.
The plucks and ticks sound distant and muffled, as if heard through a thick concrete wall.
A voiceover describes the tiled walls and the metallic ring of the room.
The plucks and ticks have a deep, rumbling echo that sounds like a large, outdoor canyon.
Model
Stage 1
Stage 2
Stage 3
Cascaded Baseline
Audex-30B
Audio-Omni
Audex-2B
UniAudio-2
Unified-IO 2
Model
S1.1
S1.2
S1.3
S2.1
S2.2
S2.3
S3.1
S3.2
S3.3
Score
Cascaded Baseline
✓
✓
✗
✓
✗
✓
✓
✓
✓
7/9
Audex-30B
✗
✓
✗
✗
✓
✗
✓
✓
✗
4/9
Audio-Omni
✓
✓
✗
✗
✗
✗
✗
✓
✗
3/9
Audex-2B
✗
✓
✗
✗
✓
✗
✓
✓
✗
4/9
UniAudio-2
✗
✗
✗
✗
✗
✗
✗
✗
✗
0/9
Unified-IO 2
✗
✗
✗
✗
✗
✗
✗
✗
✓
1/9
speechE meaning and interactionTrail directions asked again, then spelled out slowly
What to noticeInteraction and meaning. The baseline is near-perfect (9 of 9); unified models hear the words but miss how the speakers relate, bottoming out at 1 of 9.
Stage 1 - GenerateT2A
promptSteady wind through dry grass. A woman says 'Take the left fork past the fallen tree.' A man close by answers 'Got it, the left fork past the tree.'
intended audioSteady wind through dry grass. A woman says 'Take the left fork past the fallen tree.' A man close by answers 'Got it, the left fork past the tree.'
Stage 2 - EditStyle Transfer
promptThe responder changes his line to 'Hold on. Which one was it.' and nothing follows.
intended audioSteady wind through dry grass. A woman says 'Take the left fork past the fallen tree.' The man answers 'Hold on. Which one was it.' Nothing follows.
Stage 3 - Counterfactual editAdd
promptThe speaker repeats her direction at half speed, stretching and separating every syllable after the responder's question.
intended audioSteady wind through dry grass. After the man's question, the woman repeats her direction at half speed, every syllable stretched and separated.
Show all nine questions (three per stage)
Stage 1 - Generate
S1.1What environment and speakers are heard in this recording?
A windy outdoor setting with dry grass, featuring a female voice giving directions followed by a male voice repeating them in agreement.
A windy outdoor setting with dry grass, featuring a female voice giving directions followed by a male voice asking a confused question.
A transcript of a conversation about a tree path set in a windy field.
A quiet indoor room where a female voice and a male voice discuss a map without background noise.
A windy forest setting with two female voices arguing over which path to take.
A rainy street setting with a male voice giving directions to a female voice.
S1.2In what order do the sounds occur in this recording?
The sound of wind blowing is heard first, followed by a female voice speaking, and then a male voice speaking immediately after.
The male voice speaks first, followed by the female voice, over a continuous background of wind.
The female voice speaks first in silence, and then the wind starts blowing when the male voice begins speaking.
The male voice speaks, followed by a gust of wind, and then the female voice speaks.
A timeline diagram showing wind starting at second zero, female text at second two, and male text at second five.
The two voices speak simultaneously over the sound of wind.
S1.3What is the manner of the second speaker's response?
The second speaker responds in an agreeing manner, repeating the core instruction of the first speaker.
The second speaker responds in a confused manner, asking for clarification.
The second speaker responds in an angry manner, refusing to follow the instruction.
The second speaker responds in a whispered, secretive manner as if hiding from someone.
The text of the response is written in all-caps to show emphasis.
The second speaker responds with a short, non-verbal grunt of acknowledgment.
Stage 2 - Edit
S2.1What happens in this scene?
The second speaker asks a clarifying question and is met with absolute silence from the first speaker.
The second speaker repeats the direction in agreement and the scene continues normally.
The second speaker asks a question and the first speaker immediately answers it.
A script showing a question followed by a blank line representing silence.
The second speaker starts to speak but is interrupted by a sudden gust of loud wind that drowns out the voice.
The second speaker walks away without saying anything, leaving only the sound of wind.
S2.2What is the sequence of events heard in this recording?
The female voice speaks first, followed immediately by the male voice asking a question, ending in silence.
The female voice speaks, the male voice agrees, and then both continue talking.
The female voice speaks, the male voice agrees, and then the wind fades out.
A flowchart showing female speech leading to a male question block, ending in a null state.
The male voice asks a question first, followed by the female voice, ending in silence.
The female voice speaks, followed by a long silence, and then the male voice asks a question at the very end.
S2.3How does the interaction end after the second speaker finishes speaking?
The interaction ends abruptly without any further speech or response after the question.
The interaction ends with the first speaker confirming the direction again in a reassuring tone.
The interaction ends with the second speaker repeating the direction in agreement.
The audio file waveform shows a flat line immediately following the second speaker's utterance.
The interaction ends with the sound of footsteps walking away into the distance.
The interaction ends with the background wind becoming significantly louder and masking all other sounds.
Stage 3 - Counterfactual edit
S3.1Which of these is present in this scene?
The background wind, the initial female instruction, and the male speaker's question are all present before the final speech segment.
The male speaker's question is missing, and the female speaker repeats her direction immediately after her first utterance.
The initial female instruction and the male speaker's agreement are present, followed by silence.
A multi-track audio layout showing wind, female voice, male voice, and a separate slow-tempo track.
Only the background wind and a single female instruction are present, with no male speaker heard at all.
The male speaker's question is present, but there is no background wind or initial female instruction.
S3.2What is the temporal sequence of all speech events in this recording?
The female voice speaks, the male voice asks a question, and then the female voice speaks a second time at a much slower pace.
The female voice speaks, the male voice agrees, and then the female voice speaks again at normal speed.
The female voice speaks, the male voice asks a question, and then the recording ends in silence.
A timeline diagram showing three text blocks where the third block is visually stretched out.
The female voice speaks twice in a row, followed by the male voice asking a question.
The male voice asks a question, the female voice speaks, and then the female voice speaks again slowly.
S3.3What is the manner of the final speech segment in this recording?
The final utterance is delivered at half speed, with highly elongated and distinctly separated syllables.
The interaction ends in silence immediately after the male speaker's question, with no final speech segment.
The final utterance is delivered in a normal, conversational speed and tone.
The waveform display shows a stretched audio region with wider spacing at the end of the file.
The final utterance is delivered in a whispered, quiet tone as if trying not to be overheard.
The final utterance is delivered in an angry, shouting manner with high volume.
Model
Stage 1
Stage 2
Stage 3
Cascaded Baseline
Audex-30B
Audio-Omni
Audex-2B
UniAudio-2
Unified-IO 2
Model
S1.1
S1.2
S1.3
S2.1
S2.2
S2.3
S3.1
S3.2
S3.3
Score
Cascaded Baseline
✓
✓
✓
✓
✓
✓
✓
✓
✓
9/9
Audex-30B
✗
✓
✓
✗
✗
✗
✓
✓
✗
4/9
Audio-Omni
✗
✓
✓
✗
✗
✗
✓
✓
✗
4/9
Audex-2B
✗
✗
✓
✗
✗
✗
✗
✓
✗
2/9
UniAudio-2
✗
✗
✗
✗
✗
✓
✗
✗
✗
1/9
Unified-IO 2
✗
✗
✗
✗
✗
✗
✗
✗
✓
1/9
Main results
Coherence is the share of the nine questions per test answered correctly on the model's own generated audio. The Cascaded Baseline chains specialist generators and editors, read by a frontier audio judge; unified models should in principle surpass it by sharing representations across heads, so it marks what is attainable today, not a hard ceiling.
Table 1. Main results on all 432 questions (chance 16.7 percent; stages out of 144; objective metrics over 144 clips). CM = correct modality; WER vs. authored transcripts; KL / FAD / FD vs. the baseline clip set, comparable within a column only. Bold: baseline row and best unified model per column. ± is the 95% CI half-width from a cluster bootstrap over the 48 tests (5000 resamples).
Never-flagged subset (354 questions), n = 47 to 116 per cell. Bars grow from 0 with 95% CI whiskers. Physical causality (B) is where the baseline most clearly clears the models: 70.2 versus a best unified score of 48.9. Inferring what an intervention does to a scene is the hardest family for every model tested. On source composition (A) the unified models draw level with the baseline - they can hear what is present, but not what it caused.
16 tests per type by design, so every type is on a level field; Audio-Omni is the native edit arm. The cascaded-versus-unified gap is widest on sound (68.5 vs 46.0, a 22.5-point gap) and narrowest on speech (59.3 vs 54.2), consistent with the long history of speech generation and recognition. Both classes are weakest on music, which carries the lowest ceiling of any type (56.2).
Findings
Coherence is low and drops under editing.Every subject peaks at Stage 1 and falls once editing begins. Models render a change and then cannot read it back.
Render quality does not track self-coherence.Audio-Omni leads on the distributional generation metrics yet trails Audex-30B on coherence - a strong Render phase whose Evaluate phase cannot read it.
The weakness is not specific to unified models.Even the specialist baseline stops well short of reliability and generates the wrong modality roughly 40 percent of the time.
Human listening study
A blind pairwise A/B preference study over the models' generation heads. Seven raters judged 315 same-prompt comparisons (21 per model pair). Two unified models are preferred over the specialist Cascaded Baseline in generative fidelity, even though they trail it on self-coherence.
Overall win rate (%)
0
20
40
60
chance 50
46.7
Casc
54.3
A-Omni
51.4
A-30B
38.1
A-2B
31.4
UA-2
31.4
UIO2
Casc
A-Omni
A-30B
A-2B
UA-2
UIO2
Casc
47.6
42.9
38.1
66.7
38.1
A-Omni
42.9
47.6
61.9
42.9
76.2
A-30B
47.6
33.3
47.6
76.2
52.4
A-2B
33.3
23.8
28.6
38.1
66.7
UA-2
33.3
52.4
0.0
42.9
28.6
UIO2
42.9
0.0
47.6
19.0
47.6
Opponent model →
Left: overall win rate, non-tie preference share per model, dashed line marks the 50% chance level. Right: head-to-head win rate, row model versus column opponent, darker means the row model won more often. Inter-rater reliability on decisive judgments was moderate (Fleiss κ = 0.377, Krippendorff α = 0.429, 71.9% raw agreement). Audio-Omni (54.3%) and Audex-30B (51.4%) both clear the Cascaded Baseline (46.7%): generative fidelity does not track self-coherence.
Scope and honesty
TORUS is balanced by design at 16 tests per audio type. The Cascaded Baseline is not sold as a ceiling: unified models see broader data and share representations, so the expectation is that they should ultimately surpass cascades of specialists. The coupled construction gate is a diagnostic, not a calibrated filter - flagged-and-repaired questions are retained after human verification, with flags shipped as provenance. The question bank is human-verified. All tests are English-only, and findings are bounded by the five subjects and the specialist pool available at run time.
Anonymous project page for double-blind review. Links disabled during review.