Synthetic vocal bursts — six classes, listen and decide

Rare vocal-burst LoRA adapters for the LAION MOSS voice-acting model, trained on a corpus that was manufactured: a real human burst spliced into the middle of a real utterance. The machine's verdict is printed next to every clip; it is not the arbiter here.

The circularity you are being asked to check. The labels on the source clips in section A are the top-1 prediction of laion/vocalburst-classifier-single. The verdict beside them comes from laion/vocal-burst-detector-v2. Those are two heads of the same family, so training on one and scoring with the other is partly self-confirming. On the 264 clean source clips the evaluation detector names the source dataset's own label 3.4 % of the time and something in the right family 34.1 % of the time. The question for the ear is simply: is the thing labelled “Frustrated Groan” a frustrated groan?

What was measured

Old adapter versus new adapter, same prompts, same seeds, paired at the prompt level (seed*1000003 + j depends on the prompt only, so both arms drew identical sampling noise). n is the number of paired prompts, 10 per cell, three samples each, averaged within a prompt before the t — clips of one prompt are not independent. Strict and family-relaxed are shown side by side and neither is hidden.

inline cue, merge weight w = 1.0

classSTRICT — detector says the exact classFAMILY-RELAXED — a near-miss inside the family countswrong burstΔWERn
prompts
old→newΔtold→newΔtold→new (Δ)guardrail
frustrated_groan0.000→0.000+0.0000 (ctrl)0.000→0.533+0.533*+7.240.633→0.967 (+0.333)+0.05010
cackle0.000→0.033+0.033+1.000.300→0.733+0.433*+2.900.700→0.867 (+0.167)+0.03710
shriek0.000→0.000+0.0000 (ctrl)0.000→0.033+0.033+1.000.467→0.833 (+0.367)+0.06210
snicker0.000→0.000+0.0000 (ctrl)0.267→0.667+0.400*+2.570.767→0.800 (+0.033)+0.07110
cough0.000→0.000+0.0000 (ctrl)0.100→0.200+0.100+1.150.800→0.867 (+0.067)+0.11710
sniff0.000→0.033+0.033+1.000.067→0.300+0.233*+2.690.733→0.900 (+0.167)+0.04310
POOLED+0.011+1.43+0.289*+6.04+0.189+0.06360

solo cue, merge weight w = 1.0

classSTRICT — detector says the exact classFAMILY-RELAXED — a near-miss inside the family countswrong burstΔWERn
prompts
old→newΔtold→newΔtold→new (Δ)guardrail
frustrated_groan0.000→0.000+0.0000 (ctrl)0.100→0.300+0.200+1.960.500→0.833 (+0.333)+0.02710
cackle0.000→0.033+0.033+1.000.400→0.533+0.133+0.770.467→0.700 (+0.233)+0.08010
shriek0.000→0.033+0.033+1.000.000→0.133+0.133*+2.450.467→0.633 (+0.167)+0.02710
snicker0.000→0.000+0.0000 (ctrl)0.267→0.400+0.133+1.080.533→0.600 (+0.067)+0.54710
cough0.000→0.000+0.0000 (ctrl)0.000→0.133+0.133*+2.450.467→0.833 (+0.367)+0.02010
sniff0.000→0.000+0.0000 (ctrl)0.100→0.133+0.033+0.560.567→0.533 (-0.033)+0.14710
POOLED+0.011+1.43+0.128*+3.10+0.189+0.14160
Read the wrong-burst column in the same breath as the gain. Every family-relaxed gain here is bought with more wrong bursts, and on some classes the wrong rate rises further than the gain does: shriek inline gains +0.033 at family level while its wrong-burst rate rises +0.367. What the new adapters reliably do is make the model emit a burst; making it emit the right burst is only partly achieved, and the strict metric — the detector saying the exact class name — is +0.000 for five of six classes at every weight, exactly as it is for the adapters they replace.
Why strict stays at zero. The detector names the source dataset's own label 3.4 % of the time. For frustrated_groan the label the new adapter produces is Exhausted Groan (3 → 25 detections at w = 1.0) — and that is the same label the detector gives the real clean Frustrated Groans the corpus was built from. The strict metric is measuring the naming convention, not the sound.

Dose — pooled over the six classes

w = 0 is the same-model control, where every paired difference is exactly zero by construction. Compare at w ≤ 1.0: a separate dose study measured w = 1.5 as unusable on filtered buckets (pooled −0.174, t −4.70, ΔWER +0.329).

cuewΔ strictΔ familytΔ wrongΔWERn
inline0.0+0.000+0.0000 (ctrl)+0.000+0.00060
inline0.25+0.000+0.011+0.41+0.156-0.00760
inline0.5+0.000+0.050+1.84+0.244+0.01660
inline1.0+0.011+0.289*+6.04+0.189+0.06360
solo0.0+0.000+0.0000 (ctrl)+0.000+0.00060
solo0.25-0.006+0.028+0.93+0.044+0.04860
solo0.5+0.000-0.006-0.23+0.067+0.15660
solo1.0+0.011+0.128*+3.10+0.189+0.14160

The corpus

Each row is speech A → gap → burst → gap → speech B, with A and B the same speaker at genuineness ≥ 0.75, cuts placed in local RMS minima within ±150 ms of each join and 10 ms equal-power crossfades, gaps redrawn per row from 80–400 ms, and the burst levelled independently (−5…+2 dB jitter) so loudness cannot become the shortcut. EN/DE is held at exactly 50/50.

classfamrows plannedsource burstsreusespeakersEN / DEdetector-alignedgate Crows trained
frustrated_groanvbr20002308–9 (mean 8.7)16051000 / 1000138/227 (61%)97.8%1897
cacklevb20003775–6 (mean 5.31)14721000 / 1000376/403 (93%)97.7%1894
shriekvb20002159–10 (mean 9.3)15891000 / 1000215/254 (85%)98.0%1901
snickervbr15901599–11 (mean 10.0)1229795 / 795159/160 (99%)97.9%1473
coughvb830839–11 (mean 10.0)757415 / 41583/348 (24%)98.5%755
sniffvbr330339–11 (mean 10.0)330165 / 16533/186 (18%)98.8%309

frustrated_groan is the odd one out: it was built before the detector-alignment filter existed and uses the unfiltered pool of 230 bursts. Its numbers are the more conservative ones, and it is the class whose adapter was not rebuilt — replacing a measured result with an unmeasured one is not an improvement.

The five classes that are not measurable with this instrument

These were attempted and abandoned. They are not failures of the corpus: the evaluation detector has no vocabulary for them, so neither the strict nor the family-relaxed metric can score them at all.

classsource clipsaligned under the selection mapselection familyfamily under the scoring mapwhat the detector hears instead (top 3 of the source clips)
clicks_tongue2210 (0.0%)nonenone — cannot be scoredAhem 81, Breathy Giggle 54, Low Mumble 46
heavy_breathing2009 (4.5%)breathnone — cannot be scoredContented Sigh 67, Exhausted Groan 45, Low Mumble 37
soft_whistle4461 (0.2%)mouthnone — cannot be scoredLow Mumble 327, Exhausted Groan 83, Ahem 23
spitting1390 (0.0%)nonenone — cannot be scoredno_burst 24, Breathy Giggle 24, Exasperated Sigh 19
tongue_click3300 (0.0%)mouthnone — cannot be scoredChildlike Giggle 81, Breathy Giggle 76, Low Mumble 68

A. The clean source, and what voice conversion does to it

Left: the burst exactly as it comes out of laion/vocal-bursts-clean. Right: the same clip after Chatterbox VC to a real corpus speaker. One variable. Conversion failed its pre-registered gate — identity preserved 0.360 against a 0.45 threshold, family 0.470 against 0.60, speech-control word error 0.122 → 0.267 — so the shipped corpus splices the bursts unconverted, and the burst carries a different voice from the speech around it. That is the largest known weakness of this data, and this section is where you can hear why the alternative was worse.

B. Speech controls

Real utterances through the identical conversion, with the transcript on both sides. If speech survives and bursts do not, the loss belongs to the burst.

C. Whole assembled rows, all six classes

Two utterances by one speaker with a burst spliced between them, and the stated burst span beside the audio. Listen for a click at the joins and for whether the burst sounds like the same person — it is not, and that is the point of listening.