Synthetic vocal bursts — six classes, listen and decide
Rare vocal-burst LoRA adapters for the LAION MOSS voice-acting model, trained on a corpus that was manufactured: a real human burst spliced into the middle of a real utterance. The machine's verdict is printed next to every clip; it is not the arbiter here.
laion/vocalburst-classifier-single. The verdict beside them comes from
laion/vocal-burst-detector-v2. Those are two heads of the same family, so training on
one and scoring with the other is partly self-confirming. On the 264 clean source clips the
evaluation detector names the source dataset's own label 3.4 % of the time and
something in the right family 34.1 % of the time. The question for the ear is simply:
is the thing labelled “Frustrated Groan” a frustrated groan?What was measured
Old adapter versus new adapter, same prompts, same seeds, paired at the prompt
level (seed*1000003 + j depends on the prompt only, so both arms drew identical
sampling noise). n is the number of paired prompts, 10 per cell, three samples each,
averaged within a prompt before the t — clips of one prompt are not independent. Strict and
family-relaxed are shown side by side and neither is hidden.
inline cue, merge weight w = 1.0
| class | STRICT — detector says the exact class | FAMILY-RELAXED — a near-miss inside the family counts | wrong burst | ΔWER | n prompts | ||||
|---|---|---|---|---|---|---|---|---|---|
| old→new | Δ | t | old→new | Δ | t | old→new (Δ) | guardrail | ||
| frustrated_groan | 0.000→0.000 | +0.000 | 0 (ctrl) | 0.000→0.533 | +0.533* | +7.24 | 0.633→0.967 (+0.333) | +0.050 | 10 |
| cackle | 0.000→0.033 | +0.033 | +1.00 | 0.300→0.733 | +0.433* | +2.90 | 0.700→0.867 (+0.167) | +0.037 | 10 |
| shriek | 0.000→0.000 | +0.000 | 0 (ctrl) | 0.000→0.033 | +0.033 | +1.00 | 0.467→0.833 (+0.367) | +0.062 | 10 |
| snicker | 0.000→0.000 | +0.000 | 0 (ctrl) | 0.267→0.667 | +0.400* | +2.57 | 0.767→0.800 (+0.033) | +0.071 | 10 |
| cough | 0.000→0.000 | +0.000 | 0 (ctrl) | 0.100→0.200 | +0.100 | +1.15 | 0.800→0.867 (+0.067) | +0.117 | 10 |
| sniff | 0.000→0.033 | +0.033 | +1.00 | 0.067→0.300 | +0.233* | +2.69 | 0.733→0.900 (+0.167) | +0.043 | 10 |
| POOLED | — | +0.011 | +1.43 | — | +0.289* | +6.04 | +0.189 | +0.063 | 60 |
solo cue, merge weight w = 1.0
| class | STRICT — detector says the exact class | FAMILY-RELAXED — a near-miss inside the family counts | wrong burst | ΔWER | n prompts | ||||
|---|---|---|---|---|---|---|---|---|---|
| old→new | Δ | t | old→new | Δ | t | old→new (Δ) | guardrail | ||
| frustrated_groan | 0.000→0.000 | +0.000 | 0 (ctrl) | 0.100→0.300 | +0.200 | +1.96 | 0.500→0.833 (+0.333) | +0.027 | 10 |
| cackle | 0.000→0.033 | +0.033 | +1.00 | 0.400→0.533 | +0.133 | +0.77 | 0.467→0.700 (+0.233) | +0.080 | 10 |
| shriek | 0.000→0.033 | +0.033 | +1.00 | 0.000→0.133 | +0.133* | +2.45 | 0.467→0.633 (+0.167) | +0.027 | 10 |
| snicker | 0.000→0.000 | +0.000 | 0 (ctrl) | 0.267→0.400 | +0.133 | +1.08 | 0.533→0.600 (+0.067) | +0.547 | 10 |
| cough | 0.000→0.000 | +0.000 | 0 (ctrl) | 0.000→0.133 | +0.133* | +2.45 | 0.467→0.833 (+0.367) | +0.020 | 10 |
| sniff | 0.000→0.000 | +0.000 | 0 (ctrl) | 0.100→0.133 | +0.033 | +0.56 | 0.567→0.533 (-0.033) | +0.147 | 10 |
| POOLED | — | +0.011 | +1.43 | — | +0.128* | +3.10 | +0.189 | +0.141 | 60 |
shriek inline gains +0.033 at family level
while its wrong-burst rate rises +0.367. What the new adapters reliably do is make the model
emit a burst; making it emit the right burst is only partly achieved, and the
strict metric — the detector saying the exact class name — is +0.000 for five of six
classes at every weight, exactly as it is for the adapters they replace.frustrated_groan the label the new adapter produces
is Exhausted Groan (3 → 25 detections at w = 1.0) — and that is the
same label the detector gives the real clean Frustrated Groans the corpus was built from. The
strict metric is measuring the naming convention, not the sound.Dose — pooled over the six classes
w = 0 is the same-model control, where every paired difference is exactly zero by construction. Compare at w ≤ 1.0: a separate dose study measured w = 1.5 as unusable on filtered buckets (pooled −0.174, t −4.70, ΔWER +0.329).
| cue | w | Δ strict | Δ family | t | Δ wrong | ΔWER | n |
|---|---|---|---|---|---|---|---|
| inline | 0.0 | +0.000 | +0.000 | 0 (ctrl) | +0.000 | +0.000 | 60 |
| inline | 0.25 | +0.000 | +0.011 | +0.41 | +0.156 | -0.007 | 60 |
| inline | 0.5 | +0.000 | +0.050 | +1.84 | +0.244 | +0.016 | 60 |
| inline | 1.0 | +0.011 | +0.289* | +6.04 | +0.189 | +0.063 | 60 |
| solo | 0.0 | +0.000 | +0.000 | 0 (ctrl) | +0.000 | +0.000 | 60 |
| solo | 0.25 | -0.006 | +0.028 | +0.93 | +0.044 | +0.048 | 60 |
| solo | 0.5 | +0.000 | -0.006 | -0.23 | +0.067 | +0.156 | 60 |
| solo | 1.0 | +0.011 | +0.128* | +3.10 | +0.189 | +0.141 | 60 |
The corpus
Each row is speech A → gap → burst → gap → speech B, with A and B the same speaker at genuineness ≥ 0.75, cuts placed in local RMS minima within ±150 ms of each join and 10 ms equal-power crossfades, gaps redrawn per row from 80–400 ms, and the burst levelled independently (−5…+2 dB jitter) so loudness cannot become the shortcut. EN/DE is held at exactly 50/50.
| class | fam | rows planned | source bursts | reuse | speakers | EN / DE | detector-aligned | gate C | rows trained |
|---|---|---|---|---|---|---|---|---|---|
| frustrated_groan | vbr | 2000 | 230 | 8–9 (mean 8.7) | 1605 | 1000 / 1000 | 138/227 (61%) | 97.8% | 1897 |
| cackle | vb | 2000 | 377 | 5–6 (mean 5.31) | 1472 | 1000 / 1000 | 376/403 (93%) | 97.7% | 1894 |
| shriek | vb | 2000 | 215 | 9–10 (mean 9.3) | 1589 | 1000 / 1000 | 215/254 (85%) | 98.0% | 1901 |
| snicker | vbr | 1590 | 159 | 9–11 (mean 10.0) | 1229 | 795 / 795 | 159/160 (99%) | 97.9% | 1473 |
| cough | vb | 830 | 83 | 9–11 (mean 10.0) | 757 | 415 / 415 | 83/348 (24%) | 98.5% | 755 |
| sniff | vbr | 330 | 33 | 9–11 (mean 10.0) | 330 | 165 / 165 | 33/186 (18%) | 98.8% | 309 |
frustrated_groan is the odd one out: it was built before the detector-alignment filter existed and uses the unfiltered pool of 230 bursts. Its numbers are the more conservative ones, and it is the class whose adapter was not rebuilt — replacing a measured result with an unmeasured one is not an improvement.
The five classes that are not measurable with this instrument
These were attempted and abandoned. They are not failures of the corpus: the evaluation detector has no vocabulary for them, so neither the strict nor the family-relaxed metric can score them at all.
| class | source clips | aligned under the selection map | selection family | family under the scoring map | what the detector hears instead (top 3 of the source clips) |
|---|---|---|---|---|---|
| clicks_tongue | 221 | 0 (0.0%) | none | none — cannot be scored | Ahem 81, Breathy Giggle 54, Low Mumble 46 |
| heavy_breathing | 200 | 9 (4.5%) | breath | none — cannot be scored | Contented Sigh 67, Exhausted Groan 45, Low Mumble 37 |
| soft_whistle | 446 | 1 (0.2%) | mouth | none — cannot be scored | Low Mumble 327, Exhausted Groan 83, Ahem 23 |
| spitting | 139 | 0 (0.0%) | none | none — cannot be scored | no_burst 24, Breathy Giggle 24, Exasperated Sigh 19 |
| tongue_click | 330 | 0 (0.0%) | mouth | none — cannot be scored | Childlike Giggle 81, Breathy Giggle 76, Low Mumble 68 |
A. The clean source, and what voice conversion does to it
Left: the burst exactly as it comes out of laion/vocal-bursts-clean.
Right: the same clip after Chatterbox VC to a real corpus speaker. One variable. Conversion
failed its pre-registered gate — identity preserved 0.360 against a 0.45 threshold, family
0.470 against 0.60, speech-control word error 0.122 → 0.267 — so the shipped corpus
splices the bursts unconverted, and the burst carries a different voice from the speech
around it. That is the largest known weakness of this data, and this section is where you can hear
why the alternative was worse.
B. Speech controls
Real utterances through the identical conversion, with the transcript on both sides. If speech survives and bursts do not, the loss belongs to the burst.
C. Whole assembled rows, all six classes
Two utterances by one speaker with a burst spliced between them, and the stated burst span beside the audio. Listen for a click at the joins and for whether the burst sounds like the same person — it is not, and that is the point of listening.