Introduction
TapTools is a collection of Max/MSP objects with roots back to 1999, rebuilt in 2026 on a portable DSP kernel library (this repository) with thin Max wrappers (the TapTools-Max package). This book is its field guide, in the tradition of the AmbiTap and SampleRateTap books and MuTap's Quieting the Loop: one chapter per object family, written for the person patching at 11 pm, not for the person grading a DSP exam.
Each chapter makes the same promises:
- It says what the thing is for — and, near the end, when it is the wrong tool, because every tool is sometimes the wrong tool.
- Every performance claim is measured, not remembered. The numbers in these
chapters come from the kernel's own test suite and from the executed
verification notebooks in
notebooks/, which drive the same C++ code the Max objects compile through a C ABI. When a chapter says "47 dB", a notebook cell measured 47 dB, and you can re-run it. - Trade-offs are stated as trades. Knobs that buy something always pay with something; the chapters try to name both sides.
The book is organized the way a patch is:
- Part I — Sources: the virtual-analog oscillator,
tap.vco~, including its analog-character section and the honest Moog recipe. - Part II — Filters: the morphing Simper SVF (
tap.svf~), the transistor ladder (tap.ladder~), and the envelope filter modeled on the Snow White AutoWah (tap.autowah~) — with its hardware-calibration harness. - Part III — Strings, rooms, and spirals: exact true-stereo convolution
(
tap.convolve~) and the two GRM Tools recreations — the tuned comb bank (tap.5comb~) and the pitch-accumulating shimmer loop (tap.pitchaccum~). - Part IV — Tape and time: the Eno recreations — the Discreet Music
two-machine tape loop (
tap.discreet~), the Music for Airports incommensurate loop bank (tap.airport~), the generative event garden (tap.garden~), and the components they decompose into. - Part V — The machines you ride: the Radiohead family — objects whose
point is the performance surface rather than a setting. The multi-head tape
echo (
tap.tapecho~), the live buffer-stutter rig (tap.stammer~), the two-stage fuzz (tap.fuzz~), the granular scrub pad (tap.scrub~), the two Ondes Martenot diffuseurs as standalone driven resonators (tap.metallique~,tap.palme~), and the Ondes Martenot voice itself (tap.ondes~, withtap.triode~andtap.touche~). - Part VI — The spectral set: the 24-band vocoder (
tap.vocoder~), the per-bin spectral gate (tap.nr~), and the bin remapper (tap.spectra~). - Part VII — The rhythm section: the Roland recreations — the TB-303 voice,
its diode-ladder filter, and its sequencer (
tap.303~,tap.diode~,tap.303.seq~), and the eight TR-808 voice channels with their row sequencer (tap.808.*,tap.808.seq~). - Part VIII — Staying in tune: the pitch corrector (
tap.tune~) and the detection/resynthesis machinery it stands on. - Part IX — The pedalboard: the stompbox recreations — the voiced feedback
overdrive (
tap.overdrive~), chasing the TS-lineage feedback pedals rather than a waveshaping curve. - Part X — The machine, file by file: the SampleRateTap-style deep dives — one chapter per kernel header, deriving the math, reviewing the code, and recording why each algorithm is written the way it is, alternatives and all. Parts I–IX are for driving the objects; Part X is for trusting them — or changing them.
- Part XI — Recipes: whole patches chasing specific sounds — the TR-808 kits behind four decades of records, the three-oscillator Moog voice — with settings you can check against the reference pages and the honest accounting of what each ingredient buys.
More chapters land as objects mature; the utility and Jitter objects live in their reference pages, where they belong.
The oscillator and its knobs
Play a perfectly calculated sawtooth for ten seconds and you will learn
something uncomfortable: perfection sounds like a diagram. Every cycle
identical, every harmonic exactly where the textbook puts it, nothing moving —
the ear files it under test tone and stops listening. Now play three of them,
each a few cents off the others and each wandering a little, into a filter that
pushes back — and the same arithmetic becomes a synthesizer. This chapter is
about tap.vco~: what it generates, what each attribute trades, and how to get
from the diagram to the instrument — including the honest version of the Moog
recipe.
Companion material: the object's reference page (docs/tap.vco~.maxref.xml)
and help patcher (help/tap.vco~.maxhelp in the TapTools-Max package) wire up
every control in this chapter; the
verification notebook
shows every number quoted here as an executed, plotted measurement.
One phase, four shapes, no aliasing panic
Inside the object there is a single master phase ramping from 0 to 1 at the
frequency you asked for. Everything else is a way of reading that phase: a
sine reads it through sin, a saw stretches it to ±1, a pulse compares it to
the pulse width, and the triangle integrates the pulse (the classic analog
trick, reproduced digitally because it behaves so well). The continuous shape
parameter (0 sine → 1 triangle → 2 saw → 3 pulse) crossfades adjacent readings
of the same phase, so a shape sweep glides through hybrid waveforms without
resetting anything.
The digital oscillator's ancient enemy is aliasing: a naive saw's harmonics
march past Nyquist and fold back as inharmonic garbage. tap.vco~ suppresses
this with polyBLEP — each waveform discontinuity is rounded across ±1 sample by
a polynomial that closely matches what a band-limited step would do. Measured
against a naive saw at 3951 Hz (a B7, ugly on purpose): the 13th harmonic folds
back to 3.4 kHz, where the naive saw puts it at −27 dB and tap.vco~ puts it at
−74 dB — 47 dB of alias suppression right where the ear is most offended.
Two things about this are worth knowing so they don't surprise you:
- The waveforms look "not band-limited" on a scope. Expected. The BLEP correction touches two samples per edge — at 440 Hz that's 2 of ~109 samples per cycle — and there is none of the Gibbs ripple that brickwall band-limited waves show, because nothing is truncated. The shape stays essentially ideal; the spectrum is what's controlled.
- Alias suppression is not alias elimination. Push the fundamental into the kilohertz range and distant fold-backs remain, tens of dB down. For melodic and bass registers they are simply gone.
The wiring
frequency (signal or float) FM, in Hz (signal) sync (signal)
| | |
+-----+--------------------------+---------------------+-----+
| tap.vco~ |
+-----------------------------+-------------------------------+
|
(signal) the waveform
- Inlet 1 sets the frequency — a float sets the attribute, a signal drives it with true per-sample resolution.
- Inlet 2 is through-zero linear FM, calibrated in Hz: the input adds directly to the effective frequency. Drive it past the carrier and the phase genuinely runs backward (that's the "through zero" — the classic DX-style sideband sound stays coherent instead of collapsing). Measured: a 500 Hz sine carrier under ±900 Hz of FM stays bounded at exactly 1.0 peak and puts its sidebands where the textbook says.
- Inlet 3 is hard sync: every rising zero crossing of the input resets the phase, with sub-sample accuracy and an alias correction on the reset. Measured: a 187 Hz slave synced to a 110 Hz master emerges periodic at 110.1 Hz — the pitch follows the master, the timbre follows the slave's frequency, which is the whole trick of sync sweeps.
Single-channel, like every TapTools DSP object: wrap it in mc. for stacks,
and keep reading, because the analog section was designed around exactly that.
One phase, many readings — with polyBLEP correcting the edges and the analog section injecting in exactly two places.
The knobs, one by one
frequency, and gliding
Hz, from LFO rates (0.01 Hz) to 20 kHz. Every parameter in the object rides a
per-sample ramp whose length is the smooth attribute (ms, default 20) — and
on frequency that ramp is portamento. Set smooth to 60–100 ms, send note
frequencies as floats, and you have the Minimoog glide, no extra objects. For
stepped pitch, set smooth low; for per-sample modulation, use the signal
inlet (which bypasses smoothing entirely — you are the smoothing).
shape and waveform
shape is the continuous morph; the waveform sine|triangle|saw|pulse message
snaps it to a corner. The corners are the pure shapes; everything between is a
crossfade of neighbors on the shared phase. Slow shape sweeps are an
underrated modulation destination — the morph is click-free by construction
(measured: a 2-second sweep from 0 to 3 keeps its RMS within a factor of ~5 and
never drops out).
pw — pulse width
Percent, 1–99, audible as shape approaches 3. The calibration is exact: a
bipolar pulse at duty d must average 2d−1, and the measured means at 10/25/50 %
are −0.800/−0.500/+0.000. PWM by an LFO into pw (via messages, riding the
smooth ramp) is the cheapest "two oscillators" impression one oscillator can
give.
gain, presets, interp
gain is output level in dB. Sixteen preset slots store every parameter
(store 1 … store 16), and recall morphs to a slot over interp
milliseconds (or an explicit time: recall 3 4000) — every parameter riding
its ramp simultaneously, shape included. A preset morph across two very
different voicings is a patch element in its own right.
seed — which unit you own
Everything random in this oscillator — the drift walk, the jitter noise, and
(below) the component tolerances — is generated deterministically from seed.
Same seed, same render, bit for bit: your mixes reproduce and the test suite
can pin behavior exactly. Different seeds decorrelate. The mental model that
pays off: a seed is a serial number. One tap.vco~ with seed 7 is a
particular oscillator that came off the line; seed 8 is the unit next to it in
the crate. An mc. stack with per-voice seeds is a set of instruments, not
copies of one.
The analog section
Here is why a hardware oscillator sounds alive, reduced to what a DSP model can honestly act on. A real VCO is unstable at two time scales — it wanders over seconds (thermal drift) and trembles over milliseconds (noise in the core) — it is mis-calibrated in a structured way (the V/oct converter is exact at its trim point and increasingly wrong away from it), and its waveforms carry the circuit's fingerprints (a bowed ramp, a rounded reset corner, a duty cycle that isn't quite 50 %). None of these is large. All of them are always present, all slightly different from unit to unit, and the ear reads their sum as alive long before it can name any of them.
tap.vco~ models each one with its own control, all in real units, all
deterministic per seed, and all exactly zero by default — the default
object is the ideal oscillator, and the kernel's test suite pins that at
imperfect 0 every seed renders bit-identically.
drift — the slow wander (cents)
A random walk: sample-and-hold noise at ~2 Hz smoothed through a ~0.5 Hz one-pole, scaled to the depth you set. This is the thermal story — the pitch center strolling around over seconds. In a unison stack it is the difference between "chorus effect" and "three players": chorus modulation is periodic and shared; drift is aperiodic and per-voice. Ranges: 3–8 cents reads as a well-serviced vintage instrument; 15–25 as a charming one; 50+ as a broken one.
jitter — the fast tremble (cents)
New with this chapter: the short-time companion — noise at ~80 Hz through a ~40 Hz smoother, so the pitch trembles cycle-to-cycle instead of strolling. Measured at 10 cents depth: the relative spread of individual periods is 2.7×10⁻³ (a few cents, exactly as labeled), against 2×10⁻⁷ for the ideal oscillator — four orders of magnitude more micro-instability, still nothing like vibrato. This is the control that stops a sustained single oscillator from sounding frozen. Ranges: 1–4 cents is felt more than heard; 8–15 is audible grit on pure waveforms.
detune and track — the calibration story (cents, cents/octave)
detune is the static offset — the coarse fact that oscillator 2 was never
exactly oscillator 1. track is subtler and very analog: cents of error per
octave from A440, the exponential converter drifting from its trim point.
Measured with track 5: exactly 0.0 cents at A440, +15.0 cents three octaves
up, −15.0 three octaves down, a clean line through the middle. Solo it is
nearly invisible; in a stack played across the keyboard it is why vintage
unisons get wider — and slightly wilder — up the neck. Ranges: real
serviced hardware tracks within 1–3 cents/octave; ±5 is a synth that needs its
yearly appointment.
imperfect — the circuit's fingerprints (0..1)
One knob for the waveform-shape story, scaled by per-seed component tolerances so each seed misbehaves in its own direction:
- the saw ramp bows into the familiar shark-fin (a visible, scope-obvious shape change; spectrally it is mostly a phase effect — stated here so you don't chase magnitude changes that aren't there),
- the saw's reset corner rounds off: a gentle one-pole closing from ~22 kHz toward ~8 kHz — measured 6.4 dB down at the 40th harmonic (17.6 kHz) at full imperfection; extreme top-end air, traded for warmth,
- the triangle goes asymmetric, and this one is very audible in the spectrum:
the ideal triangle's 2nd harmonic sits at −185 dB (i.e., absent); at
imperfect 0.8it rises to −34 dB relative to the fundamental — even harmonics, the classic "warm" giveaway, - the sine picks up mild waveshaper color, and the pulse width takes a small static offset (so two "50 %" pulses from two seeds beat against each other the way two real units do),
- the whole unit takes a static pitch offset of up to a couple of cents.
Ranges: 0.2–0.4 is a healthy vintage unit; 0.6–0.8 is character you can point to in a mix; 1.0 is a unit with a story. At 0, every seed is the same ideal machine — the analog section never costs you the reference oscillator.
The performance section
Where the analog section models what the circuit does on its own, these controls model what a hand does to it — added after the Recipes chapters had to teach a scaling formula to get constant-width vibrato out of the Hz-calibrated FM inlet.
vibrato/vibrato_rate— a sine LFO on the pitch, depth in cents (0–100) and rate in Hz, so ten cents is ten cents in every register. Measured: at a commanded ±100 cents the peak cycle-to-cycle deviation reads 90–110 cents, and the modulation crosses its mean at exactly twice the commanded rate (pinned by test).vibrato_delay— the singing control: the vibrato fades in through a one-pole with this time constant (ms), re-armed on every new note (every frequency-target change), so held notes bloom and passing notes stay plain. Pinned: early deviation under 60 % of settled, and shallow again right after a note change. The signal-rate frequency inlet deliberately does not re-arm — there, you are the modulation.bend— pitch bend in semitones (±24), riding the standardsmoothramp: the wheel, as an attribute. Pinned within 5 cents of the commanded interval.
All of it is deterministic with no randomness — and at depth 0 the output is bit-identical to the ideal oscillator (pinned), so the reference instrument is still free.
The Moog recipe, honestly
The sound everyone wants from this object is three oscillators into a ladder. Here is the recipe, with the honest accounting of which ingredient does what. Rendered A/B demos of exactly this patch (through the real kernels) live in the notebook material.
| voice | frequency | detune | drift | seed |
|---|---|---|---|---|
| 1 | f | −4 c | 8 c | 11 |
| 2 | f | +5 c | 8 c | 22 |
| 3 | f ÷ 2 | +2 c | 10 c | 33 |
- All three:
@shape 2(saw),@smooth 70for glide,@jitter 3,@track 2,@imperfect 0.3. - Sum them (scale by ~1/2.8) and feed
tap.ladder~:@mode lp24 @resonance 0.35 @drive 9 @asym 0.45 @comp 0.25.
What each ingredient buys, in order of importance:
- The stack itself. Three free-running voices at ±cents is most of the
sound. The beating between them is the fatness; the octave-down third voice
is the weight. (
tap.vco~free-runs like hardware — no per-note phase reset — so the beat pattern is different on every note. That, not any single voice's tone, is the big analog tell.) - The ladder.
driveinto the tanh stages compresses and colors the stack;asymadds the even harmonics of mismatched transistors;compkept low preserves the authentic passband droop as resonance rises. A perfect saw into a driven asymmetric ladder sounds more "Moog" than an imperfect saw into a clean one — spend your character budget here first. - Glide. 60–100 ms of
smoothon the note changes. Iconic, and free. - The analog section. Drift keeps the beating from ever repeating; jitter un-freezes sustains; per-seed tolerances make the three voices three units. This is seasoning — essential in the way salt is, invisible in the way salt is.
Omit in reverse order when CPU or taste says so.
When it is not the right tool
- You need an exact test signal. Actually — it is the right tool:
imperfect 0(the default) is the mathematically ideal oscillator, and the test suite holds it there. Just don't reach for the analog section and a measurement mic in the same patch. - You want evolving spectra from one voice — wavetables, granular motion,
additive drift. This oscillator's spectrum is fixed per shape by design;
morph
shapeor FM it, but a wavetable oscillator is a different instrument. - You want chorus. Twenty cents of drift on one voice is not a chorus; it is a seasick oscillator. Chorus is a delay effect — use one.
- Noise. The bottom of the
shaperange is a sine, not a noise source;tap.noise~has five colors of the real thing.
Checkpoint
One master phase, four shapes and their hybrids, polyBLEP keeping the folded
harmonics ~47 dB down. Frequency glides on smooth, FM is in honest Hz and
survives through zero, sync locks pitch to the master while timbre stays yours.
The analog section is four controls with real units — slow drift, fast
jitter, structured mis-calibration in track, circuit fingerprints in
imperfect — all scaled by per-seed component tolerances, all exactly off by
default, all deterministic: a seed is a serial number. And the Moog recipe is
mostly the stack and the ladder — let the oscillator's imperfections season,
not carry.
The filter that morphs
Every synthesizer needs one filter it can trust with anything: a bass line, a
noise sweep, a parametric EQ move, an audio-rate modulation stunt. tap.svf~
is that filter — a state-variable design in Andy Simper's trapezoidal
(zero-delay-feedback) formulation, the same lineage as the filters in Ableton
Live, including Auto Filter's Morph type. This chapter is what each attribute
trades, and why the design earns the trust.
Companion material: the reference page and help patcher in the TapTools-Max package, and the verification notebook, where every number below is an executed, plotted measurement.
Why "state-variable," and why this one
A state-variable filter computes all its responses — lowpass, bandpass,
highpass, notch — from the same two internal states at once, which is what
makes continuous morphing between them possible at all. The classic digital
version (Chamberlin) famously misbehaves at high cutoffs and under fast
modulation. Simper's TPT formulation fixes both: the tuning is prewarped
(exact all the way to Nyquist) and the filter is unconditionally stable
under per-sample cutoff modulation — the property that later let this same
kernel become the sweep engine inside tap.autowah~. The notebook slams the
cutoff across five octaves with a 90 Hz LFO under full-band noise; the output
stays bounded, no oversampling tricks required.
Two integrators in a zero-delay loop; every response — and the morph — is three multiplies downstream of the same two states.
The knobs, one by one
type — the discrete responses, the morph, and the EQ family
Ten responses from one core. The classics — lowpass, highpass, bandpass, notch, peak, allpass — plus:
morph: one continuous parameter sweeps LP → BP → HP → notch → LP (0 → 0.25 → 0.5 → 0.75 → 1). The corners are bit-identical to the discrete modes — measured max difference exactly 0 — so morphing to a corner is that filter. A slow morph under a held chord is a patch element the discrete modes can't give you.bell,lowshelf,highshelf: the parametric-EQ trio from Simper's coefficient tables, with a ±24 dBgain. Measured: a +12 dB bell peaks at +12.00 dB; a −9 dB low shelf lands −9.00 dB in its plateau and 0.00 dB on the other side. These always run a single 2nd-order section — cascading would square the boost, soorderis ignored for them, on purpose.
order — 2, 4, or 8 poles that stay flat
Orders 2/4/8 (12/24/48 dB per octave) run as a cascade with the Butterworth Q spread, so at resonance 0 the response is maximally flat and sits at −3.01 dB at the cutoff regardless of order — measured −3.01 at every one, with slopes of 12.3/24.7/49.4 dB per octave. The trade against a naive cascade of identical sections (which droops long before fc): none. This is just the correct way to stack poles.
resonance — normalized, and honest about the top
0 to 1: 0 is the Butterworth-flat base, 1 is the edge of self-oscillation.
Resonance sharpens only the final section of a cascade, so you get one
clean resonant peak on a flat passband instead of a compounding stack of
peaks. A q message converts to and from engineering Q if you think in those
units.
circuit — clean or driven
cleanis the pure linear filter: cheapest, transparent, never oversampled. Also the reference: the EQ modes and every measured Bode plot above are this circuit.drivenaddsdrive(dB) into a tanh limiter on each section's band node — an OTA-flavored color stage, oversampled (1/2/4×, default 2×). Two measured consequences: a 200 Hz tone through +18 dB of drive grows odd harmonics that simply do not exist in the clean circuit (the 3rd harmonic appears out of the numerical floor, ~140 dB up), and at resonance 1.0 the filter self-oscillates at the cutoff — measured 999.7 Hz for a 1 kHz setting, amplitude bounded by the saturator. It needs a ping to start: a perfectly silent filter is a fixed point.
frequency, the right inlet, and smooth
Float or attribute sets the cutoff through the anti-zipper ramp (smooth,
ms). A signal in the right inlet takes over per sample — that's the path
for audio-rate filter FM and for envelope-follower patches. Sixteen preset
slots morph via store/recall over interp milliseconds, everything
gliding together.
Recipes
- The synth voice:
@type lowpass @order 4 @resonance 0.4, envelope into the frequency inlet. Order 4 is the "synth filter" slope; order 2 is the polite one; order 8 is a wall. - The DJ sweep:
@type morph, sweepmorph0 → 0.5 while easingfrequency— the LP-through-BP-to-HP arc is the whole move in one parameter. - Tone control:
bell/shelves with modest gains. It measures exact, so trust the numbers you type. - A sine with character:
@circuit driven @resonance 1, ping it, and tune withfrequency— a self-oscillating test-tone-with-a-temper.
When it is not the right tool
- You want the classic squelchy 4-pole growl. That's a transistor-ladder
sound — resonance that compresses the passband, saturation inside the
loop. Next chapter:
tap.ladder~. - You need many static EQ bands. One
tap.svf~per band works, but a dedicated multiband EQ (ortap.filter~, the RBJ multimode biquad) is the boring, correct choice. - You want the filter to follow your playing. That's
tap.autowah~, which is this filter plus an envelope detector and a sweep law.
Checkpoint
One TPT core, every response as an output mix: discrete modes, a morph whose corners are bit-identical to them, and an exact parametric-EQ trio. Butterworth-spread orders stay −3.01 dB flat at any slope; resonance sharpens only the last section; the driven circuit adds tanh color and true bounded self-oscillation at the cutoff. Unconditionally stable under per-sample modulation — which is why other objects build on it.
The transistor ladder
Some filters are tools; this one is a character actor. The four-stage
transistor ladder — the Moog circuit — colors everything it touches: the
resonance pushes back against the bass, the stages saturate into one another,
and at the top of the resonance range it stops filtering and starts singing.
tap.ladder~ is a zero-delay-feedback model of that circuit with a tanh
saturator in every stage. This chapter is what each control trades, and what
the measurements say the model actually delivers.
Companion material: the reference page and help patcher in the TapTools-Max
package, and the verification notebook —
every number below is an executed measurement. For the linear ladder — the
cheap, polite Stilson/Smith model — see tap.fourpole~; this object is its
nonlinear sibling.
What the model gets right
Two things separate a serious ladder model from a filter with a "Moog" label:
- Tuning that survives the top octaves. The classic digital shortcut goes audibly flat as the cutoff rises. This model is prewarped ZDF: measured self-oscillation lands at 1000.2 Hz for a 1 kHz cutoff (0.02 % error) — and, the part that's actually hard, 8009 Hz for an 8 kHz cutoff (0.11 %). You can play the resonance like an oscillator anywhere on the keyboard.
- Nonlinearity inside the loop, not bolted on. Each stage saturates, and the feedback fights the saturation the way the hardware does. That is where the compression, the "sag," and the bounded self-oscillation come from.
The whole filter: four stages, one loop. The red tap sets resonance, the amber paths are the comp bargain and the Xpander mode taps.
The knobs, one by one
frequency and the right inlet
Cutoff in Hz; a signal in the right inlet drives it with true per-sample
resolution. Like everything here it rides the smooth ramp when set by
message.
resonance — up to and past the edge
0 to 1.1. At 1.0 the loop gain reaches the oscillation threshold; above it
the filter sings at the cutoff, amplitude-limited by the tanh stages (ping it
to start — silence is a fixed point). Under the edge, resonance does the
authentic ladder thing: it eats your passband (see comp).
drive — how hard to lean on the stages
Input gain (dB) into the saturating ladder. Measured THD on a 100 Hz tone:
0.5 % at 0 dB, 3.5 % at 8, 16.5 % at 16, 33 % at 24 — a smooth walk from
"slightly thick" to "fuzz pedal's cousin." All odd harmonics, because tanh is
symmetric — which is exactly why asym exists.
asym — the even harmonics of real hardware
Real transistors don't match; their operating points sit slightly off-center,
and that asymmetry is where a hardware ladder's even-harmonic warmth lives.
asym (0..1) models the mismatch. Measured on a driven tone: the 2nd
harmonic sits at −156 dB (numerically absent) at asym 0 and rises to
−18.6 dB relative to the fundamental at 0.6. One honest warning from the
reference page: an asymmetric saturator can produce slight signal-dependent
DC — follow with tap.dcblock~ if something downstream cares.
comp — the passband bargain
A real ladder trades passband level for resonance: the feedback subtracts
from the input. Measured at resonance 0.9: the passband sits at −13.2 dB with
comp 0 (the authentic droop) and at 0.0 dB with comp 1 (fully restored).
Vintage behavior or modern behavior — your call, continuously.
mode — pole mixing, the Xpander trick
lp24, lp12, bp12, bp24, hp12, hp24: mixing the ladder's stage taps yields
whole families of responses from the same four poles (the Oberheim Xpander's
famous trick). Measured small-signal slopes: 23.4 dB/oct for lp24, 11.7 for
lp12. The resonance and saturation behavior carries into every mode — a
resonant bp24 through drive is a very different animal from tap.svf~'s
clean bandpass.
oversample — paying for the saturation honestly
The tanh stages generate harmonics past Nyquist that fold back as inharmonic alias tones. Measured on a hard-driven 5 kHz tone: going from 1× to 4× oversampling drops the non-harmonic (alias) energy by 13.5 dB. The default 2× is the working compromise; use 4× when you drive high notes hard, 1× when you're filtering bass and counting CPU.
solver — fast or exact
The nonlinear loop can be solved with one predictor-corrector pass (fast,
the default) or by Newton iteration to convergence (exact,
circuit-simulation accuracy). They are audibly identical until drive and
resonance are both pushed hard; exact is there for when you want to know,
and for renders where CPU is free.
Recipes
- The bass patch:
tap.vco~saw stack (see the oscillator chapter's Moog recipe) →@mode lp24 @resonance 0.35 @drive 9 @asym 0.45 @comp 0.25. Keepcomplow; the droop is the vintage glue. - The acid line:
@resonance 0.85 @drive 15, envelope into the frequency inlet, and let the resonance fight the saturation. - The kick synthesizer:
@resonance 1.05, ping it with a click, and ridefrequencydown fast — a self-oscillating ladder is a sine with attitude.
When it is not the right tool
- Transparent filtering. Every pole of this filter has an opinion. For
surgical work use
tap.svf~(clean circuit) ortap.filter~. - Morphing responses. The pole-mix modes switch; they don't glide.
Continuous response morphing is
tap.svf~'smorph. - CPU-constrained patches that just need "4-pole lowpass."
tap.fourpole~is the linear ladder at a fraction of the cost — no saturation, no oversampling, no opinions.
Checkpoint
A prewarped ZDF four-stage ladder with tanh in every stage: self-oscillation
in tune within 0.11 % even at 8 kHz, drive that walks THD from 0.5 % to 33 %,
asym switching on the even harmonics of mismatched transistors, comp
choosing between authentic passband droop and modern flatness, pole-mixed
multimode outputs, and oversampling that measurably pays down the
saturation's aliasing. The character filter — spend your tone budget here.
The pedal that listens
A wah pedal is a filter with a foot attached. An auto-wah cuts out the
foot: it listens to how hard you play and sweeps the filter for you — hit a
string and the filter opens; let it ring and the filter settles back down.
tap.autowah~ models a specific, beloved instance of the idea: the Mad
Professor Snow White AutoWah, Björn Juhl's OTA-based envelope filter,
grounded in the traced circuit and the published behavior. This chapter is
how to drive it — and how we will know, measurably, when the model matches
the pedal.
Companion material: the reference page and help patcher in the TapTools-Max
package; the design document (plans/tap.autowah~.md in TapTools-Max) with
the full hardware research; and the
validation notebook,
which measures everything below and ends with a cell waiting for recordings
of the real pedal.
What the hardware is, in one paragraph
A 2-pole state-variable filter (an LM13700 OTA circuit) whose frequency is pushed up from a resting point by an envelope detector — a diode and a capacitor, charged fast, discharged at a rate you set. Four knobs: Sensitivity (how hard your signal drives the sweep), Decay (how fast it falls back), Bias (the resting frequency), Resonance (the Q). Published sweep: 250 Hz to about 2.5 kHz — a throaty, vocal range, deliberately unlike the quack of a Mu-Tron-style filter. One secret feature: with Sensitivity at minimum it becomes a fixed, manually swept filter — the "cocked wah."
The model composes tap.svf~'s Simper core (one 2nd-order section, driven
per sample — the modulation stability that filter chapter promised, cashed
in) behind a rectifier → attack/release follower and an exponential sweep
law. The measured control behavior:
- Sweep law: cutoff = bias · 2^(sweep · range). Measured against the design curve across the full envelope range: max error 0.000 cents. The law lives in one function on purpose — if the real pedal turns out to sweep linearly in Hz, one function changes and nothing else moves.
- Timing: attack set to 2 ms measures 1.94 ms; decay set to 250 ms measures 256 ms, and the release fits a pure exponential with residual σ = 0.004 — an RC discharge, like the hardware.
A detector, a law, and a borrowed filter — the amber chain is everything this object adds to the SVF it composes.
The knobs, one by one
sensitivity — the trigger level, and the secret mode
Detector input gain in dB (−60..+24). Tune it to your instrument and touch:
too low and only your hardest hits open the filter; too high and everything
pins at the ceiling (a tanh soft knee compresses hard playing into the top
rather than slamming a rail). At −60 the envelope is exactly off and the
object becomes the cocked wah: a fixed resonant filter with bias as the
manual sweep control. Factory preset 4 ships that voicing.
decay — the personality knob
How fast the filter falls back to bias, in ms (10..5000). Fast (tens of
ms) gives a wah articulation on every note — the funk setting. Slow
(hundreds of ms up) gives classic auto-wah swells that ride your phrasing.
This is the knob to perform.
bias and range — where the sweep lives
bias is the resting frequency (default 250 Hz, the hardware's home);
range is the sweep span in octaves above it (default 3.3, the hardware's
250 → ~2500 Hz). Both go far beyond the hardware if you want them to, and
direction 1 sweeps down from bias instead — a TapTools extension the
pedal never had.
resonance, mode, drive
resonance (0..1) is the Q — the vocalness. mode picks the filter tap:
lowpass is the stock voicing; bandpass is the circuit's other node, a known
hardware mod — quackier and noticeably quieter. drive (dB) engages the
saturating SVF circuit for OTA-flavored color; 0 keeps it pure.
attack, mix, and the rectifier
The hardware's attack is fast and fixed; ours defaults to the same 2 ms but
is exposed (0.05..100 ms) for softer onsets. mix is an equal-power dry/wet
the pedal never had — 100 % (wet-only) is the hardware. And under the hood
the detector's rectifier is selectable: full-wave (default, cleaner
tracking) or half-wave (the traced single-diode topology). Measured: the
half-wave detector carries 2.6 % signal-rate ripple on a low tone against
the full-wave's 0.7 % — a real, quantified flavor difference awaiting the
hardware A/B.
The sidechain inlet and the envelope outlet
A signal in the right inlet takes over the detector: one sound wahs another (a kick opening a pad is the classic). The right outlet emits the envelope (0..1) as a signal — the detector as a free modulation source for anything else in the patch.
Factory voicings
Preset slots 1–4 ship guitar, bass (lower bias, tighter range — the
GB pedal's instrument switch, as a preset you can morph to), slow swell,
and cocked wah. recall 2 4000 morphing from guitar to bass over four
seconds is its own effect.
How we'll know it matches
The validation harness is built and proven on ground truth. An STFT
peak-trajectory extractor recovers the swept resonant peak from wet audio
alone — no dry reference needed — and against the kernel's own cutoff trace
it correlates at 0.979 in log-frequency, with the small measured offset
(−37 cents) close to what resonant-peak physics predicts (−20 cents at that
Q). The same code runs on demo videos and on the real pedal. A Snow White is
on order; when it arrives, reamped recordings drop into
notebooks/reference/ and the notebook's last cell overlays hardware against
model. Disagreements map one-to-one onto kernel constants. That pass may flip
the default filter tap or the sweep law — both are flagged, isolated, and
waiting.
When it is not the right tool
- You want the filter on a knob or LFO, not your dynamics. That's
tap.svf~with a signal in its frequency inlet — this object's own core, without the detector. - Your source has no dynamics. A static pad through an auto-wah is a static filter. Feed the sidechain something rhythmic instead.
- You want the Mu-Tron quack. Different circuit, different voicing —
raise
resonance, trymode 1, but know you're modding a Snow White, not summoning a Mu-Tron.
Checkpoint
An envelope detector with a fast attack and a musician's decay knob, driving
a 2-pole resonant filter up from bias through an exponential law that
measures exact to the design. Sensitivity at the floor is the cocked wah;
the sidechain inlet and envelope outlet make the detector patchable; the
factory slots hold the four voicings that matter. And the model doesn't ask
to be trusted — the extractor that will judge it against the real pedal is
already built, already proven, and already waiting in the notebook.
Borrowed rooms
Every room you have ever heard is a filter: clap once, and what comes back —
the impulse response — is everything the room will ever do to any sound.
Convolution reverb plays your signal through that recording. tap.convolve~
does it in true stereo against an impulse response held in a buffer~, using
the standard engine of the genre (uniformly-partitioned overlap-save FFT
convolution), and its defining property is worth stating up front: it is
exact. Not "high quality" — exact. This chapter is what that buys, what it
costs, and how to drive the two knobs that actually matter.
Companion material: the reference page and help patcher in the TapTools-Max package, and the verification notebook — the first notebook in this repo, and the template for all the others.
Exact, measured
The engine splits the IR into blocksize-sample partitions, transforms each
once, and multiply-accumulates in the frequency domain over a delay line of
past input spectra. The bookkeeping is intricate; the result is not. Measured
against a direct time-domain convolution of the same float32 IR: maximum
difference 3×10⁻¹² — double-precision noise. Change the block size and
the output doesn't change either (64/256/1024 agree within 2×10⁻¹²,
latency-removed). An impulse through the engine reconstructs the IR to
5.5×10⁻¹⁴, and a synthetic 0.60 s-RT60 reverb measures back at 0.599 s. There
is no "character" in this engine to audition; the character is entirely in
the IR you load.
The wiring, and the one real cost
Stereo in, stereo out, IR from a named buffer~. The cost: latency of
exactly blocksize samples — verified for every block size — on top of
your I/O latency. That is the entire quality/latency dial:
blocksizesmall (64–128): tight enough for live input; more CPU per sample (more, smaller FFT batches).blocksizelarge (512–2048): cheapest; latency grows to match. For a send/return reverb on a mix bus, nobody hears 21 ms of pre-delay you didn't ask for — except you, so usepredelaydeliberately instead.
maxsize reserves capacity (both are locked in while DSP runs and apply on
restart). The rest of the surface mirrors tap.verb~ so the two reverbs read
as siblings: mix, gain, predelay, normalize (energy-based, so quiet
and hot IRs land at comparable levels), bypass, mute.
The wiring: one FFT in, one IFFT out, and a multiply-accumulate that is the only cost growing with IR length.
True stereo, by channel count
A stereo room isn't two mono rooms: sound from the left source arrives at the
right ear too. The engine runs the full 2×2 matrix — LL, LR, RL, RR — and the
buffer~'s channel count selects the topology:
- 4+ channels: true stereo, all four paths (measured: a signal sent only left emerges on the right at exactly the cross-feed path's gain, 0.600 expected, 0.600 measured, with zero leakage where paths are silent).
- 2 channels: dual mono — L and R convolved separately, no cross-feed.
- 1 channel: the same mono room on both sides.
Loading rooms while the music plays
IRs are analysed off the audio thread and published atomically into a double-buffered slot: swapping IRs mid-performance neither clicks nor drops (measured RMS across the swap instant: 21.9 before, 22.1 just after), and one block later the output is bit-identical to an engine that had the new IR from the start. Load rooms like presets; the engine doesn't flinch.
Recipes
- The honest room: a measured IR (church, plate, spring — the internet is
full of them),
@mix 25 @normalize 1, and resist the urge to EQ the IR itself before tryingpredelay— 10–30 ms of it buys clarity for free. - Not a reverb at all: an IR is any filter. A single click is a delay; a strummed guitar body is a body simulator; a vowel is a formant filter. Convolution doesn't know it's supposed to make reverb.
- True-stereo width: record or synthesize the four paths with a genuinely different LR/RL from LL/RR — the cross-feed is where "being in the room" lives.
When it is not the right tool
- You want to design the reverb — decay knobs, damping, modulation,
gated tails. A static IR can't do any of that;
tap.verb~(the algorithmic Moorer reverb, with its own oversampling and limiter) can. - Zero-latency insert on a live path. The engine costs
blocksizesamples, full stop. At 64 that's 1.3 ms — small, not zero. - Time-varying convolution. Swaps are click-free but discrete; the engine doesn't interpolate between rooms.
Checkpoint
Partitioned convolution is exact linear convolution — measured to 10⁻¹² —
with one honest cost, blocksize samples of latency, and one honest dial,
block size against CPU. True stereo comes from the buffer's channel count;
IR swaps are atomic and dropout-free; and everything the effect sounds like
is the impulse response you feed it. Borrow better rooms.
Five strings, no guitar
Feed a comb filter its own output and it stops being an EQ curiosity and
becomes a string: a resonator with a pitch, a ring time, and a temperament.
Five of them, tuned by hand, is an instrument — that was the insight of the
GRM Tools Classic "Comb Filters" plugin, and tap.5comb~ is its recreation:
five resonant combs with per-voice tuning, masters that play the whole bank,
and the preset-morph engine that made the GRM tools feel alive. This chapter
is how to tune, ring, and morph it.
Companion material: the reference page and help patcher in the TapTools-Max
package, and the grm_comb_render tool in the kernel repo, which renders the
listening-check scenarios outside Max.
What a resonant comb actually is
A delay of 1/f seconds fed back on itself resonates at f and all its harmonics — pluck it with noise and it rings like a string tuned to f. Two implementation details decide whether five of them sound like a chorus of strings or like a broken flanger, and both were the reasons this object was recreated rather than ported:
- Fractional delays. At 48 kHz, a 440 Hz comb needs a delay of 109.09 samples. Round it to 109 and the comb plays 440.37 Hz — every voice lands on a slightly wrong, slightly different wrong pitch, and the beating between voices (the whole point of a bank) is gone. The delays here are Hermite-interpolated: the tuning is continuous, and sweeps glide instead of zippering.
- No clipper in the loop. The feedback path uses a DC blocker and a precise feedback cap, not a hard limiter — high resonance rings clean instead of distorting.
One voice of five. The red ring is the string; the amber tap is the pluck position.
The knobs, one by one
freq1..5 and freq — the tuning
Per-voice frequencies (5 Hz floor, the GRM's own) — or the notes
message, which tunes up to five combs from MIDI note numbers in one list
(fractional allowed) — plus a master multiplier
(0..2) that transposes the whole bank — the master is the performance
control, gliding every voice proportionally so chords stay chords.
res1..5 and res — ring time, not feedback
Resonance is mapped to ring time on a log curve, 20 ms to 100 s, and the feedback coefficient is derived from the current delay — so a voice keeps its ring time as its pitch sweeps, instead of ringing longer at low notes and choking at high ones (the raw-feedback behavior of naive combs, and of the legacy abstraction). 50 is a decaying pluck; 80+ sustains; near 100 it is a drone that outlives your patience.
lp1..5 and lp — the string's brightness
A one-pole lowpass inside each feedback loop: every pass around the loop gets darker, which is exactly how real strings decay (highs first). Open it for metallic; close it toward a few hundred Hz for felt and thump.
warp — stiff strings
New to this recreation: a negative-coefficient allpass in the loop disperses the partials — upper harmonics round-trip faster and stretch sharp, the inharmonicity of a stiff piano string. The main tap is compensated at each voice's fundamental, so the pitch stays put while the timbre goes piano-ish, then bell-ish. At extreme warp × high tuning the loop can't get shorter than the dispersion and the pitch flattens — physical, and documented.
phase — where you pluck the string
Also new: a half-loop pickup tap. At 100 the even harmonics cancel — the sound of plucking a string exactly at its midpoint. Neutral at 0.
gain, mix, and the morph engine
Equal-power dry/wet and output gain, plus the GRM hallmark: sixteen preset
slots with timed interpolation. store 1, retune everything, store 2,
then recall 1 8000 — every frequency, resonance, and damping glides for
eight seconds through territory you never explicitly tuned. Grabbing one
parameter mid-morph overrides just that parameter. The morph is not a
transition between sounds; it is the sound.
Recipes
- The resonator chord: tune
freq1..5to a voicing (say 80/120/160/200/ 102 Hz — the legacy factory preset),reshigh, and feed it drums or speech. The input is now an excitation signal for your chord. - The piano that isn't: moderate resonance,
warp 40,lparound 3 kHz, and pluck with clicks. - The eight-second gesture: two stored extremes and a long
recall— the classic GRM move. Automate nothing else.
When it is not the right tool
- One comb, precise and plain:
tap.comb~is the single, cheaper unit. - Echoes rather than pitch: delays long enough to hear as repeats are
tap.delay~/tap.multitap~territory — a comb is a delay, but this one is tuned and normalized for resonance, not slapback. - Faithful nostalgia: this deliberately is not the legacy
tap.5comb~abstraction — the integer delays, linear feedback, and in-loop clipper it had are exactly what was retired, and the deviations are flagged in the reference page for the audition.
Checkpoint
Five Hermite-tuned resonant combs with ring time on a log map (20 ms–100 s),
per-loop damping, and two ways to bend the string physics (warp for
stiffness, phase for pluck position) — under masters that transpose the
bank and a sixteen-slot morph engine that turns retuning into performance.
Strings, chords, drones, and gestures; no guitar required.
The spiral staircase
Most pitch shifters are a one-way trip: in, transposed, out. The GRM Tools
"PitchAccum" closed the loop — the transposed signal is delayed and fed back
into the transposer, so every pass around the loop shifts it again. Set
+7 semitones and a note becomes a rising spiral: +7, then +14, then +21, each
echo climbing, the whole thing dissolving upward like light on water. That
loop is the effect everyone now calls shimmer, years before the name.
tap.pitchaccum~ is the recreation: two independent transposer-delay loops
("shadows") with the accumulation wired in. This chapter is how to climb.
Companion material: the reference page and help patcher in the TapTools-Max
package, and the grm_pitchaccum_render tool in the kernel repo for
listening checks outside Max.
The loop, and why it doesn't collapse
Each shadow is: granular transposer (±24 semitones) → delay (up to 3 s) → feedback → back into the transposer. Three design choices keep the spiral musical instead of muddy:
- Constant-level grains. The transposer sweeps two taps half a cycle apart, each windowed so the pair sums exactly to 1 at every phase and every crossfade width. The original tt_shift engine's window pair didn't quite sum flat, which imposed an amplitude ripple at the grain rate — after ten trips around a feedback loop, ripple compounds into tremolo. Here the tenth pass is as steady as the first.
- Hermite-interpolated taps. Fractional delays keep each pass in tune, so the spiral's steps are the interval you set, not the interval plus drift.
- A capped, DC-blocked loop. Feedback tops out at 0.99 with a DC blocker in the path — the spiral can run for a very long time, but it is unconditionally bounded (unit-tested at the cap).
The signature is measurable: set +7 semitones and the kernel test finds energy at +7 and +14 — the second pass, the accumulation itself.
The topology is the effect. Feedback re-enters upstream of the taps, so the staircase climbs; in the ordinary patch every echo is the same interval.
The knobs, one by one (per shadow, ×2)
trans1 / trans2 — the step of the staircase
±24 semitones, continuous. Musical intervals (+7, +12, +5) make harmony; small offsets (±0.1–0.3 st) make lush detune-echo instead of a spiral; negative values descend into the dark version nobody expects.
delay1 / delay2 — the tread depth
Up to 3 s per shadow. Short (50–150 ms) blurs the passes into a texture; long (0.5–2 s) articulates each step of the climb as an audible echo.
fb1 / fb2 — how many steps
How much survives each trip, 0–99. 30 gives two or three audible generations; 70 a long climb; 90+ a texture that essentially sustains until the transposition walks it out of range (energy shifted past the audible band is the spiral's natural exit).
xfade — the grain crossfade
GRM's Cross-fade control: the width of the grain envelope's flanks. Narrow is more articulate and more grain-rate flavored; wide is smoother and softer in attack. Because the envelope pair always sums to 1, this changes texture, never level.
The modulation section, and follow
A global LFO (with modphase offsetting shadow 2, so the two loops breathe
against each other) plus per-voice deterministic random transposition
modulation — a little of either keeps long spirals from sounding cloned.
follow (off by default) engages a pitch follower — decimated normalized
autocorrelation, confidence-gated, deliberately picking the smallest
plausible lag so it doesn't lock onto subharmonics — which adapts the grain
window toward the detected period: cleaner transposition on monophonic
sources, ignored gracefully on noise.
Presets
The sixteen-slot morph engine, as everywhere in the GRM pair: two stored
spirals and a timed recall between them is a gesture in itself.
Recipes
- Shimmer, the classic: shadow 1 at +12, delay ~400 ms,
fb1 75; shadow 2 at +7, delay ~650 ms,fb2 60; both into a reverb (tap.convolve~with a long church, ortap.verb~). The reverb is load-bearing — shimmer is spiral plus wash. The full patch has its own recipe in Part IX. - The descent: −5 and −12, long delays, moderate feedback — a staircase into the basement, much rarer and much creepier.
- Micro-thickener: ±0.15 st, 60/90 ms delays, feedback 50,
xfadewide — not a spiral at all, just an expensive-sounding widener.
When it is not the right tool
- One clean transposition, no loop:
tap.shift~is the plain shifter — same modernized engine, none of the plumbing. - Formant-true vocal shifting: granular transposition shifts formants
with the pitch; chipmunks live this way.
tap.harmony~is the dedicated tool — formant-preserving voices at fixed intervals, chords included. - Rhythmically exact multi-tap echoes: the delays here serve the loop;
tap.multitap~serves the grid.
Checkpoint
Two transposer-delay loops where the feedback re-enters the transposer, so pitch accumulates pass after pass — +7 becomes +14 becomes +21. Constant-sum grain envelopes keep the tenth pass as steady as the first; Hermite taps keep it in tune; the capped, DC-blocked loop keeps it bounded forever. Intervals are the architecture, delay is the pacing, feedback is the height — and the morph engine turns the whole staircase into something you can bend mid-climb.
The tape that forgets slowly
Every other delay in this house is kept honest by a cap: feedback stops just
short of one, because a loop that gains nothing and loses nothing will pile
up until it clips. tap.discreet~ is built on the opposite bargain. Its
regeneration goes all the way to 1.0 — legally, cleanly, forever — because
the loop forgets: every pass through the tape comes back a little darker
and a little softer than it went in. The memory loss is not a defect the
kernel tolerates; it is the mechanism that keeps the machine stable. You are
not patching a delay effect. You are renting a machine whose memory is the
instrument.
The rig it recreates is printed on the back cover of Discreet Music (Obscure/EG, 1975): Brian Eno's synthesizer feeding one Revox tape machine, the tape spooling for seconds across the room to a second machine, and the second machine's playback both sent to the speakers and folded back into the first machine's record head. It is the same two-machine system Robert Fripp ran for the No Pussyfooting loops. The tape path itself — the fractional read, the periodic wow and flutter, the in-loop coloration — follows the published tape-echo modeling literature (Arnardóttir, Abel, and Smith's AES model of the Echoplex, and Välimäki et al.'s tape-echo work). The schematic is the score; this kernel is a faithful performance of it.
Companion material: the executed notebook discreet.ipynb, which measured
every claim below, and the eno_render tool, whose discreet_basic and
discreet_sustain scenarios are the listening copies. The Max wrapper lands
in the TapTools-Max package alongside the rest of the family.
Two machines and a spool of tape; the red return is where the forgetting — and therefore the stability — lives.
loop — the tape span
loop_seconds is the distance between the machines: how long a phrase
travels before it returns. The kernel test pins the grid to the sample — an
impulse comes back at exactly one loop, bit-for-bit the first time, and
every later return lands within a sample of its grid point.
Changing the loop while audio runs is a tape-speed change, not a menu option: the read head physically glides to its new distance, and gliding a read head is doppler. Move from 0.5 s to 0.75 s over half a second and the playback drops an octave while the transport re-spools, then re-locks on pitch — the test measures 220 Hz mid-glide and 440 Hz within five cents after. There is no crossfading "digital" mode, on purpose. If a pitch bend on loop changes would ruin the patch, this is the wrong delay (see below).
regen — and why 1.0 is legal here
regen is the return level into the record head, and unlike tap.delay~'s
feedback (capped at 0.99), it reaches exactly 1.0. The notebook plays a
one-second noise burst into the loop at regen 1.0 and lets it run for twenty
seconds: the level settles and stays — no growth, no collapse — because the
wear path bounds it. The saturator's output can never exceed 1/drive
regardless of what the loop accumulates, the DC blocker keeps offsets from
stacking, and the darkening lowpass decides what survives: lows sustain,
highs surrender. The pinned scenario is blunt about the contract — it
asserts non-growth, never decay, because at regen 1.0 sustain is the
promise. Bring regen down, or darken harder, to end a piece; clear is
the eject button, and regen-1.0 material is gone for good.
darken and drive — the wear
darken_hz is the record/playback corner: every pass through the loop runs
through a one-pole lowpass at this frequency, so a bright phrase sheds its
treble generation by generation while its body lingers. This is measured,
not vibes: with the corner at 2 kHz, a 6 kHz tone loses to 0.292 of itself
per pass and a 300 Hz tone keeps 0.890 — and both numbers match the analytic
transfer of the wear path to three decimals in the executed notebook.
Generation loss, measured against regen · |H_wear|. The tape forgets treble first.
drive is the record-head saturation — the guarantee. At any drive above
zero the loop is absolutely bounded no matter the settings; at drive 0 the
path is exactly linear (a real bit-for-bit passthrough, not "almost") and
the loop leans on darkening alone. Drive around 0.5 is the tape sound;
drive high is the loop slowly compressing itself into a wash.
wow and flutter — the transport
Two sines, slow-deep and fast-shallow, breathing the play head's position. The pitch math is honest and checkable: depth times 2π times rate is the peak deviation, so 2 ms of wow at 0.5 Hz predicts ±10.9 cents — and the notebook's YIN pitch track measures 10.9. The transport is periodic and deterministic by design (no stochastic capstan drift): two renders of the same settings are bit-identical, which is also a pinned test. Set both depths to 0 for a perfectly still machine.
input_level — the performance move
The fader Eno actually rode was not the output — it was the send. Play a
few phrases into the machine, then bring input_level to zero: the loop
keeps unrolling everything it holds, worn a shade further every pass, and
the piece continues without you. That gesture — set up a system, feed it,
step away — is the whole record, and it is one setter here. mix is the
ordinary equal-power dry/wet with bitwise-exact endpoints.
Recipes
- The Discreet Music bed:
@loop 5. @regen 0.95 @darken 3500 @drive 0.4 @mix 60. Play sparse, slow phrases; stop; listen to what the tape decides to keep. - Frippertronics:
@loop 6.5 @regen 1. @drive 0.7 @darken 2200 @mix 100. Solo over yourself from a minute ago. The wash never clips and never ends until you end it. - Haunted slapback:
@loop 0.15 @regen 0.85 @wow 4. 0.9 @flutter 0.15 12.— a short loop with a seasick transport; the doppler and the wear turn a slap delay into a memory of one. - The exit: whatever is running, ride
@regenfrom 1. to 0.7 over a minute. The piece performs its own fade, oldest material first.
When it is not the right tool
- Rhythmic delays. Loop changes bend pitch by design, and there is no
tempo sync.
tap.delay~is the clean line;tap.multitap~is the pattern. - Anything that must not color the repeats. Wear is always in the loop (drive 0 removes only the saturation, not the darkening you set). If the tenth echo must equal the first, this machine is philosophically opposed.
- Loops that should line up with other loops. One machine, one spool.
For a bank of independent free-running loops, the next chapter's
tap.airport~is the instrument.
Checkpoint
Seconds of tape between two machines; a worn return path — darken, saturate,
DC-block — instead of a feedback cap; regeneration to exactly 1.0 because
forgetting is the stabilizer. Loop moves are honest tape-speed doppler, the
transport is two deterministic sines measured in cents, and the send fader
is the performance. Every number above lives twice: as an executed cell in
discreet.ipynb and as a pinned scenario in tests/discreet_test.cpp,
which CI runs on every push.
Loops that never line up
Take seven tape loops of deliberately awkward lengths — none a multiple of
another — put one soft phrase on each, and let them all turn at once. Each
loop is trivial: it plays the same thing forever. The system is not: the
phrases drift against each other, meet, part, and meet again differently,
and the pattern of coincidences does not repeat within a human afternoon.
That is "2/1" from Brian Eno's Music for Airports (Ambient 1, EG, 1978),
as he described the rig in the album's liner notes and in A Year with
Swollen Appendices: the lengths are the score, and the machine's whole job
is to keep the loops turning without an opinion. tap.airport~ is that
machine — up to eight free-running loops, each with a single head that both
plays and records, summed to stereo.
The discipline that makes it the instrument it is: nothing resets a phase. Not recording, not a level move, not a pan, not even a length change. The free-run is the composition, and the kernel treats the heads as sacred; the test suite literally hammers every setter mid-run and then checks that the heads have advanced by exactly the samples processed.
Companion material: the executed notebook airport.ipynb, which measured
every claim below, and the eno_render tool's airport_two_one scenario —
three stereo minutes of seven loops, the listening copy. The Max wrapper
lands in the TapTools-Max package alongside the rest of the family.
One loop of eight. The head plays, then records, then advances; nobody ever tells it where to be.
Record and return
record(loop, 1) punches the input onto that loop's tape at wherever its
head happens to be — there is no downbeat, no quantized punch-in, because
Eno's rig had none. Recording replaces (each phrase was recorded once, not
overdubbed), and playback reads just ahead of the write, so while recording
you hear the previous generation under the head. record(loop, 0) freezes
the tape, and freezes it bit-exactly: the pinned test compares two whole
passes of a frozen loop and requires them identical to the bit. A loop is
not a degrading medium here — it replays the same magnetic imprint every
revolution, which is why this kernel deliberately has no per-pass
generation loss (that is tap.discreet~'s physics, not a loop's).
The lengths are the score
length_seconds per loop is where the composing happens. Two loops of
24000 and 30000 samples realign only at their least common multiple —
120000 samples, 2.5 seconds — and the kernel will tell you:
composite_period_seconds reports exactly 2.5 for that pair, confirmed in
the notebook by rendering the coincidence raster and watching it repeat at
2.5 s and at no shorter lag.
Two awkward lengths and their coincidences. Stretch the lengths and the composite period leaves the room.
Then stretch toward the piece: give seven loops airport-scale lengths in awkward ratios and the composite period overflows a 64-bit sample count — the kernel reports infinity, which is not a failure mode. It is the point.
Changing a length while running is a splice: the tape keeps its content and the head re-wraps modulo the new length — never rewinding — exactly as cutting a physical loop shorter would land you mid-phrase. It can click. Splices do.
Level, pan, shade
Each loop has a slewed linear level, an equal-power pan with exact
endpoints (a hard-panned loop is bitwise absent from the far bus — the
same law as tap.multitap~), and a darken corner that shades that loop's
playback tone. The shade is a static one-pole per loop, not wear: measured
in the notebook, a 6 kHz phrase through a 1 kHz shade lands at 0.169 of its
transparent twin, against an analytic prediction of 0.169. At the band
ceiling — the default — the shade stage is bypassed entirely and playback
is bit-transparent, which is what makes the freeze and hard-pan promises
testable as bitwise facts rather than tolerances.
There is deliberately no wow here: the phasing engine of "2/1" is the
incommensurate lengths, not pitch drift. If a loop's source should breathe
like tape, run it through tap.discreet~ on the way in.
Recipes
- The terminal: seven loops,
@lengths 17.8 19.1 21.3 23.9 26.2 28.7 30.9, one sustained tone phrase recorded onto each, levels around 0.45, pans spread wide, a 4 kHz shade on two of them. Let it run. Come back in an hour; it will not have repeated. - Phase study: two loops, lengths in a near ratio (say 8.0 and 8.1), the same short phrase on both, panned hard left and right — the Reich-adjacent version, where the drift itself is the melody.
- Sound-on-sound sketchpad: one loop,
@lengths 12., record gate on a footswitch. Punch in fragments as they occur to you; the head's indifference to your downbeat is the charm. - Breathing loops: patch sources through
tap.discreet~(gentle wow, regen 0) before the record gate — tape transport on the way in, stable free-run once captured.
The same machine, in pieces
There was never a loop bank doing loop-bank things in here — there is an
array of eight identical lanes and a summing loop. That lane is now an
object of its own, tap.reel~, and three of them summed are a
tap.airport~ bitwise (pinned in tests/airport_test.cpp). Patch it
instead of using this object when you want an insert on one loop, a
varispeed on one reel, more than eight loops, or tape you actually use —
the bank buys all eight worst-case reels at DSP start regardless. See
The same machine, in pieces.
When it is not the right tool
- Synchronized looping. This machine never lines up by design. A beat-locked looper wants a phase reset on the downbeat, which is the one thing this kernel refuses to do.
- Degrading loops. A frozen loop here is bit-eternal. For material that
should wear out as it circulates,
tap.discreet~is the machine with the forgetting built in. - Dense delay textures. Eight long loops is a composition system, not
an echo;
tap.multitap~does a hundred taps without ceremony.
Checkpoint
Up to eight free-running loops, one sacred head each: record replaces at
wherever the head is, freeze is bitwise, splices re-wrap and never rewind,
and no setter touches a phase. Level, exact-endpoint pan, and a bypassable
playback shade place the phrases; the lengths do the composing, and
composite_period_seconds tells you how long until the piece repeats —
ideally, longer than you will be alive. Every number above lives twice: as
an executed cell in airport.ipynb and as a pinned scenario in
tests/airport_test.cpp, which CI runs on every push.
The garden that plays itself
The first two chapters of this part recirculate sound: tape that forgets, loops that never agree. This one recirculates decisions. Plant a note and it comes back every pass of the loop a step quieter and a step purer, until it fades below hearing and retires. Plant several and they braid. Stop planting altogether and, after a patient interval, the garden starts planting for itself — always on the scale, never in a hurry. You do not play this instrument so much as tend it, which is exactly the posture Eno kept asking for: the composer as gardener, not architect. The kernel is named for that metaphor.
What it recreates is the principle behind Brian Eno and Peter Chilvers'
generative apps (Bloom, 2008), as described in their published interviews
and in Eno's 1996 "Generative Music" talk: touch becomes note, note repeats
and fades, scale makes wrong notes impossible, idleness hands the piece to
the system. The principle only — no scale tables, timings, or sounds are
taken from the app, and its name is a live trademark of Opal Limited, which
is why this object is a garden and not a bloom. (As with tap.tune~'s
history paragraph, none of this is legal advice; the project's ship-gate is
a freedom-to-operate review.)
Companion material: the executed notebook garden.ipynb, which measured
every claim below, and the eno_render tool's garden_played and
garden_idle scenarios, the listening copies. The Max wrapper lands in the
TapTools-Max package alongside the rest of the family.
Events on a loop instead of audio on a tape — the same recirculation, one level of abstraction up.
Plant and return
note(pitch, velocity) plants: the pitch snaps to the current root and
scale at entry, a small wind chime is struck on the next sample,
and the event takes a seat at the loop's current position. Every pass, it
fires again at velocity × decay, and below floor it retires. The
notebook's staircase is the whole contract in one figure: a plant at 0.8
with decay 0.5 returns with its fundamental at exactly half the last, four
times over (measured ratios 0.500, 0.500, 0.500, 0.500), then silence, and
active_events reads zero. The whole strike fades a shade faster than its
fundamental — quieter returns are also duller, because strike hardness
couples brightness to velocity.
The return staircase: decay 0.5, floor 0.05, and a bloom that knows when it is finished.
That arithmetic is also the stability story. The family's inversion —
degradation as the stabilizer — reaches its third form here: a bloom lives
exactly ceil(log(floor/velocity) / log(decay)) passes, so the population
of live events converges by construction no matter how fast you plant.
And beneath the arithmetic sits a hard bound: sixteen chimes in a fixed
pool, the quietest stolen when a seventeenth is needed, its envelopes
re-aimed rather than reset so a steal glides instead of clicking.
soften — returns get purer, not just quieter
The chime is four decaying mode doublets at the transverse ratios of the
chosen material — by default 1 : 2.756 : 5.404 : 8.933, the free-free-tube
physics in Fletcher & Rossing — with the upper modes softer, steeper in brightness
(b, b², b³), and dying roughly as f² faster, so the fourth mode is the
few-millisecond tick of clapper contact and every strike rings down to its
fundamental. Each mode pair is split a few cents, the way a real tube's
degenerate modes are, so the tail beats slowly instead of decaying like a
lab sine. Each pass multiplies the event's brightness by soften, and
brightness is the upper modes' level: a bloom collapses toward its
fundamental as it recedes, losing its tick first — the tape chapters'
generation loss, restated in modes instead of passbands. The notebook
measures the mode-two-to-fundamental ratio shrinking by exactly soften
every single return, and the pinned tests hold each piece separately: the
tick confined to the contact, the tail's beat dipping and returning, soft
strikes duller than hard ones, high tubes ringing shorter than low.
The rack: material, flaws, and seats
material swaps what the tubes are made of — a mode, not a fader. At 0 the
rack is wind chimes, the free-free tube's 1 : 2.756 : 5.404 : 8.933; at 1 it
is a tuned bar, the mallet instrument's double-octave 1 : 4 : 10 : 20 (both
tables from Fletcher & Rossing). The table is read at strike time, so every
live bloom re-voices at its next return: the notebook measures the second
partial's energy moving cleanly from 2.756× to 4× when the material flips.
And the tube is the identity. Each pitch is a physical tube whose
imperfections are properties of the tube, not the strike: its upper modes
sit a fixed few cents off the ideal ratios (bounded by ±3 cents — the
fundamental stays true, because a maker tunes the fundamental and the
overtones land where the metal puts them), and it hangs at a fixed seat on
the stereo rack, width set by spread (0 collapses to center mono, bitwise
identical busses). Both draws come from a stateless hash of the pitch, so
the rack is the same rack in every instance and every return of a bloom
rings from the same place with the same flaws — the notebook's seat chart
is a bar per pitch, and the seed triad below is untouched because no
generator is ever consumed for it.
The scale contract
root and scale (chromatic, major, minor, and both pentatonics — plain
public-domain scale theory) define where plants may land, and quantization
happens at entry: the notebook plants all thirteen chromatic pitches from
60 to 72 into a C major-pentatonic garden and the YIN oracle reads every
sounded note on {C, D, E, G, A}. Wrong notes are not discouraged; they are
unrepresentable, which is most of why instruments in this family feel
effortless to strangers. Because quantization is at entry, changing the
scale re-pitches nothing already planted — the field changes for future
seeds only.
The gardener
idle_seconds is the patience: that long after your last plant, the wind
picks up. The gardener strikes on a calm/gust cycle — gust at 0 is a
still day, single strikes spaced about one per pass; raise it and strikes
arrive in flurries of up to five neighboring tubes within a fraction of a
second, with longer calms between, the average rate holding. The
randomness is the family's
seeded xorshift64* with the full tr808 contract, pinned as a triad: same
seed, bit-identical garden; different seed, a different garden; gardener
disabled (idle_seconds 0), the seed cannot matter at all, because the
generator is never consumed. This is the library's first randomized event
source — step_seq.h proudly promises "no randomness anywhere" — and the
seed contract is what lets a generative instrument live in a test suite
that demands reproducibility.
Recipes
- The lobby: defaults,
@idle 30. @level 0.4, plant four or five notes, walk away. The garden holds the room indefinitely, bounded. - The music box:
@decay 0.5 @soften 0.7 @idle 0 @bell 0.005 0.8 1.— no gardener, fast decay: each phrase you play unwinds itself to silence in a few passes, a wind-up toy running down. - The endless install:
@scale minorpentatonic @root 2 @idle 3. @gust 0.6 @seed 2008 @level 0.35, never touch it again. Same seed next year, same garden — gusts and all. - The still day:
@gust 0 @idle 10.— no flurries, one unhurried strike at a time, the original music-box gardener. - The marimba loft:
@material 1 @spread 1. @decay 0.7 @soften 0.8— tuned bars instead of tubes, the rack thrown wide: drier, woodier blooms that each speak from their own place in the image. - Duet:
@idle 6.and stay at the keyboard — every silence longer than six seconds, the gardener answers you; every plant of yours resets its patience.
The same machine, in pieces
The four machines inside this one — the entry quantizer, the event ring,
the chime rack, the seeded gardener — are objects too: tap.scale,
tap.bloom, tap.chime~, tap.gardener. Chained, they are this object
bitwise, gardener and all (pinned in tests/garden_test.cpp). The one
worth reaching for on its own is tap.bloom: separated from the chime it
recirculates notes and has no opinion about what sounds them, so the
principle will drive a sampler or MIDI out just as happily. One difference
to know before you patch it — out here the ring runs on Max's scheduler
rather than the audio clock, so returns land within a millisecond of the
grid instead of exactly on it. See
The same machine, in pieces.
When it is not the right tool
- Melodies with wrong notes in them. Quantization is always on;
chromatic passing tones survive only in
@scale chromatic, and micro-tonal pitches not at all. This is a fence, and it is the product. - Rhythm. Events return on the loop grid, exactly, forever — no swing,
no humanization. For patterns as rhythm,
tap.808.seq~is the machine. - Any other timbre. Two materials, one chime family, on purpose. It is an instrument, not a polysynth; for synthesis as a playground, patch oscillators.
- A stereo panner. The image is a rack of fixed seats keyed by pitch — there is no per-strike pan and no motion. For placement as a parameter, pan the object's output.
Checkpoint
Notes become events; events recirculate on a loop, quieter by decay and
purer by soften each pass, retiring below floor; a sixteen-chime pool
bounds the sound and a sixty-four-seat ring bounds the score, oldest bloom
yielding first. Two materials share the rack, every tube keeps its own
flaws and its own stereo seat, the scale makes wrong notes unrepresentable,
and a seeded gardener keeps the piece alive exactly as long as you neglect
it. Every
number above lives twice: as an executed cell in garden.ipynb and as a
pinned scenario in tests/garden_test.cpp, which CI runs on every push.
The same machine, in pieces
The two chapters before this one describe instruments you switch on and
walk away from. tap.airport~ turns seven loops; tap.garden~ tends
itself. That is the right shape for what they do, and neither is going
anywhere.
But both were monoliths by accident rather than by design. Open
airport.h and there was never a loop bank doing loop-bank things — there
was an array of eight identical lanes and a summing loop. Open garden.h
and there was a quantizer, an event ring, a chime rack, and a seeded
gardener, wired together by a class that did nothing else. The parts were
already there. Nothing outside the monolith could reach one.
So they were promoted. The lanes and the parts are objects now, and the block diagrams at the top of the last two chapters are patchable:
| Object | What it is | Was |
|---|---|---|
tap.reel~ | one free-running tape loop | a lane of tap.airport~ |
tap.chime~ | the sixteen-bell wind-chime rack | the voice pool of tap.garden~ |
tap.chime.voices~ | the same rack, one bell per outlet | — |
tap.bloom | the event ring — plant, return, fade, retire | the recirculation of tap.garden~ |
tap.scale | snap a pitch to a root and scale | the entry quantizer |
tap.gardener | the idle wind, seeded | the self-seeding half |
tap.period | when a set of loops realigns | the bank's period message |
The monoliths remain exactly what they were. This is additive: the same kernel classes, reached two ways.
"The patch is the object" is a measurement, not a slogan
It would be easy to say that three tap.reel~ summed are a
tap.airport~ and leave it there. The house rule is that claims of that
kind get measured, so this one is pinned in CI like any performance
number.
The scenario "standalone lanes summed are the bank, bitwise" in
tests/airport_test.cpp configures a three-lane bank and three standalone
lanes identically — incommensurate lengths, both exact pan endpoints and
one interior pan, one shaded darken corner and one bypassed — drives both
through the same staggered punch-in schedule for two seconds, and requires
the two stereo outputs to be equal to the bit, not to a tolerance.
Nudging one lane's level by 1e-12 fails it.
The garden's version, "the bed is exactly its components wired
together, bitwise" in tests/garden_test.cpp, does the same across
twenty seconds with the seeded gardener running — which puts the order of
random draws under test too, since that is the part a careless split moves
without anyone noticing.
Bitwise is available here because the objects' own promises are already bitwise: transparent playback of a frozen loop, exact pan endpoints, a darken stage that is genuinely bypassed at the band ceiling. A decomposition can be held to the same standard the object is.
What you get for patching it
For the airport, four things the monolith cannot give you:
- An insert on one loop. A filter, a reverse, a
tap.discreet~for tape breath on one phrase and not the others. Inside the bank every loop gets the same treatment, which is to say none. - A varispeed on one reel — the bank has one shared clock by construction.
- More than eight loops. Eight was a number, not a principle.
- Tape you actually use. The bank buys all eight worst-case reels at
DSP start whether you use them or not: about 92 MB of double tape at
the 30-second default. Three
tap.reel~buy three, about 11 MB each.
The rack has a second form worth knowing about. tap.chime.voices~ is the
same sixteen bells with each one on its own outlet, carrying its tube dry —
before the seat in the stereo image. Filter one voice and you are filtering
whichever bell happens to be in that slot, not the rack; it is a different
instrument, and there is no way to ask tap.chime~ for it. Because the pool
reassigns bells as it steals, a slot is not a pitch, so the object will tell
you which tube it is holding and what seat it would have been given. Sum the
sixteen back through those seats and you have tap.chime~ again, bitwise —
pinned by "the per-voice taps summed through their seats are the stereo
rack", across twenty strikes, four more than the pool holds, so stealing is
under test too.
It is a separate object rather than a switch because outlet count is fixed
when a Min object is built, and it is sixteen discrete outlets rather than one
multichannel outlet because min-api's mc support is inlet-side only: it sets
Z_MC_INLETS and offers no multichanneloutputs, which is what Max requires
before an external may declare a variable-channel mc outlet. That is a
limitation of the wrapper we have, not of the idea.
For the garden, the interesting one is tap.bloom. Separated from the
chime it turns out to be the most portable idea in the family, because it
recirculates notes and has no opinion about what sounds them. Point it
at makenote, at a sampler, at MIDI out, and Eno's principle — a touch
becomes a note, the note returns a little quieter each pass until it is
gone — drives an instrument that has nothing to do with wind chimes.
Splitting also made two promises directly testable that were previously
only reachable through audio. The ring's arithmetic is now countable with
no envelope tail in the way: "the ring's convergence theorem is exact
when nothing sounds it" checks four different velocity/decay/floor
triples against ceil(log(floor/velocity)/log(decay)) exactly. And the
rack's allocator can be watched directly — "the rack fills idle bells
first, then steals the quietest" fills the pool, strikes a seventeenth
tube, and measures that the faint tube lost its partial while a loud one
kept its own.
Where the seams show
Three honest costs, none of them hidden.
The garden's patch is not sample-accurate. tap.bloom and
tap.gardener run on Max's scheduler rather than the audio clock, so a
return lands within an @interval tick — a millisecond by default —
instead of exactly on the sample. Inside tap.garden~ the same ring is
sample-accurate. At loop lengths measured in seconds nobody will hear the
difference, but it is a difference, and it is why the garden's null test
lives in the kernel where both sides can share one clock, and why there is
deliberately no in-Max null test for it. Asserting a null that cannot hold
would be worse than not asserting one.
Voice stealing had to stay in the kernel. The obvious Max answer to a
sixteen-voice rack is one voice in a poly~. That answer is wrong twice:
poly~ steals round-robin, which loses the whole point — this rack steals
the quietest bell and re-aims it, so its phases keep free-running and its
seat glides rather than clicking — and poly~ does not exist off Max,
while the kernel is meant to run anywhere. So tap.chime~ is the whole
rack, and its polyphony is its own.
A bell reads silent until it has been processed once. The allocator asks each bell for its level, and a bell that has been struck but not yet processed still reports zero. Strikes issued in the same sample therefore land on the same voice instead of spreading across the pool. Inside the bed this only happens when two blooms share a loop position; it is pre-existing behaviour, and it is documented in the rack scenario rather than fixed, because fixing it would change how the object sounds.
The one thing the airport decomposition looked like it would lose is
composite_period — the report of when the whole system realigns, which
needs every length at once and so has nowhere to live inside a single reel.
That arithmetic came out of the bank as a free function instead, and
tap.period is it: hand it the lengths and it answers in seconds, inf
included. The detail that makes it trustworthy rather than merely
convenient is that it shares the reel's seconds-to-samples quantization
rather than copying it. The lcm is over sample counts, and lengths that
look commensurate written down are usually nothing of the kind once
rounded to samples — 0.5 and 0.625 seconds realign at 2.5 s, while the
terminal recipe's seven lengths leave the 64-bit range entirely. Both are
pinned in tests/airport_test.cpp.
Checkpoint
Seven objects, almost no new DSP: the same kernel classes the monoliths
hold, given names and inlets. Three tap.reel~ summed are a tap.airport~ bitwise;
tap.gardener into tap.scale into tap.bloom into tap.chime~ is a
tap.garden~ bitwise, gardener and all. Both identities are pinned
scenarios in tests/airport_test.cpp and tests/garden_test.cpp, which
CI runs on every push, and the airport's is checked again against the real
externals loaded in Max by
runtime-tests/patchers/tap.reel~-is-airport.maxtest.maxpat. The
monoliths still do what they did — every scenario that pinned them before
the split passes unchanged after it.
Four heads and a motor
The last chapter's tap.discreet~ is a machine you set up and walk away
from. tap.tapecho~ is one you keep your hands on. Same spool of tape, same
worn return path, same family — but where the Eno objects are systems that
run without you, this one is an instrument, and every parameter on it is a
hand on the machine. That is the thread through this part of the book: these
are the objects you ride.
What it recreates is the tape echo of the Copicat / Space Echo school: one
record head, a span of moving tape, several playback heads at fixed
positions along it, and a path from the heads back to the record head. Ed
O'Brien's Copicat is the reason it is here. It is a recreation of the
topology, not a circuit model of any one unit — the tape path itself is
the same published tape-echo modeling literature tap.discreet~ already
stands on (Arnardóttir, Abel, and Smith's AES model of the Echoplex, and
Välimäki et al.'s tape-echo work), and no head spacing, filter curve, or
trim value in this object is claimed as measured from a real machine.
Companion material: the executed notebook tapecho.ipynb, which measured
every number below, and the radiohead_render tool, whose tapecho_heads,
tapecho_three_head, tapecho_selfosc, and tapecho_varispeed scenarios
are the listening copies — all four performed, with the controls moving
while they render, because static settings tell you almost nothing about
this object.
span — the motor
span is the delay of a head sitting at the far end of the tape path, and
every other head sits at span times its own ratio. So span is not "the
delay time" of one echo; it is the motor speed, and moving it moves the
whole layout together.
One impulse, four heads. The returns land exactly on span × ratio.
Moving the motor while audio runs is a tape-speed change, which means it
bends pitch on the way — the same doppler contract as tap.discreet~, for
the same reason: the heads are physically moving relative to the tape.
smooth sets how long the motor takes to change speed, and therefore how
deep the bend is. There is no crossfading "digital" mode. If a pitch bend on
a delay-time change would ruin the patch, reach for tap.delay~.
heads, ratios, levels, pans — the layout
Four heads by default, evenly spaced at 0.25, 0.5, 0.75 and 1.0 of the span. That spacing is nominal — chosen because it is neutral and audibly a tape echo — and every ratio is freely settable underneath, which is how you build a three-head Copicat-style layout:
heads 3, ratios 0.333 0.667 1.
levels is per-head gain and pans places each head in the stereo field
(equal-power, with exact endpoints: a hard-panned head is bitwise absent
from the far bus). One thing to know: a head's level is also its send into
the regeneration path, as the head selector on the real machines is. Turn
a head down and you are turning down both what you hear from it and what it
feeds back.
regen, drive, darken — past unity, on purpose
Here is where this object parts company with everything else in the house.
tap.delay~ caps feedback at 0.99 so the loop is always contractive.
tap.discreet~ reaches exactly 1.0 because the wear path is the stabilizer.
tap.tapecho~ goes past 1.0 — up to 1.5 — into deliberate
sound-on-sound self-oscillation, the howl you reach for this machine to get.
It stays bounded because the saturator does. drive is record-head
saturation, and its output can never exceed 1/drive no matter what the loop
accumulates, so the tape is bounded by the input plus regen/drive whatever
the loop gain. The measurement is the point:
Regeneration at 1.4 — well past unity — plateaus under the saturator's ceiling at every drive.
Because that bound only exists while the saturator is engaged, the
effective regeneration is capped back to 1.0 whenever drive is 0 — and the
cap is applied per sample, so dropping drive mid-howl lands the loop rather
than letting it run away. The attribute keeps its value and takes effect
again when drive returns. Twelve seconds of ring at regen 1.4 measures a
growth ratio of 1.007 between the two late windows: it plateaus, it does not
climb.
darken is the per-pass corner. Every trip through the regeneration path
runs through a one-pole lowpass, so the repeats lose treble generation by
generation — measured at 0.2915 of a 6 kHz tone per pass against 0.2920
predicted, and 0.8895 of a 300 Hz tone against 0.8898. Riding darken
while the loop howls is a performance control, not a set-up step; it is
what turns a howl into a swell and back.
wow and flutter — one motor, one path
The transport is the family's deterministic pair of sines, and one motor moves the whole tape path, so a speed error displaces every head together. The pitch math is checkable in closed form: depth times 2π times rate is the peak deviation, so 2 ms at 0.5 Hz predicts ±10.88 cents and the notebook's pitch track measures 10.91. Two renders of the same settings are bit-identical — periodic and deterministic by design, with stochastic capstan drift a documented non-goal, because bit-exact renders are what let the oracle test exist at all. Set both depths to 0 for a still machine.
The one that is not a knob
With the tape path neutralized — no transport error, no regeneration — a
one-head echo is bitwise tap.multitap~ with one tap. Same Hermite
read, same fractional position, same equal-power pan law. That is not a
curiosity; it is the whole design claim, measured: this object is
composition over the shared tape machinery rather than a second
implementation of it, and tape_loop.h needed no changes at all to serve a
topology it was not written for. The appendix has the derivation.
Recipes
- The Copicat:
heads 3, ratios 0.333 0.667 1.with@span 390 @regen 0.6 @drive 0.9 @darken 2600 @wow 0.9 0.9 @mix 50. Heads down the middle, a tired transport, repeats that thicken as they recirculate. - A wide slap: four heads,
pans -0.7 0.5 -0.35 0.8,@span 480 @regen 0.45 @drive 0.4 @mix 45. The layout does the widening; no chorus needed. - Sound-on-sound:
@drive 0.7 @regen 1.35, then bring@inputto 0 and take your hands off. Ride@darkendown to 1400 while it howls, then@regen 0.55to bring it home.clearis the emergency stop. - The dive:
@smooth 3000, then@span 200→@span 900. Three seconds of tape slowing down, with everything already on the tape bending with it.
When it is not the right tool
- Tempo-locked delays. Span changes bend pitch by design and there is no
sync.
tap.delay~is the clean line. - A wash you set and leave. That is
tap.discreet~, one chapter back — same machinery, opposite posture. - Independent free-running loops. One motor moves every head here. For
loops that drift against each other,
tap.airport~. - Clean repeats. Wear is always in the regeneration path;
drive 0removes the saturation, not the darkening.
Checkpoint
A motor and up to four heads along one tape path; the motor moves them
together and bends pitch doing it. Regeneration goes past unity into
self-oscillation, bounded by the saturator rather than a gain cap, and
capped back to 1.0 the moment drive leaves. The transport is two
deterministic sines measured in cents. And with the tape path neutral the
whole object collapses, bitwise, into a delay this library already had —
which is how you know it is composition and not a rewrite. Every number
above lives twice: as an executed cell in tapecho.ipynb and as a pinned
scenario in tests/tapecho_test.cpp, which CI runs on every push.
The part that comes apart
There is a moment at the end of "Go To Sleep" where the guitar stops being a
guitar. It does not fade, it does not filter — it starts eating itself,
firing fragments of the bar you just heard in an order nobody played. That
sound came out of a Max patch Jonny Greenwood built and performs live. So
tap.stammer~ has an odd position in this package: it is a Max stutter
object, in a Max package, for a technique that was invented in Max.
None of which means anything was copied. This is an original design in the brassage tradition (Roads, Microsound) — a continuously recorded buffer, a rhythmic grid, and dice. What the band's rig contributes is the knowledge of what the object is for, which turns out to be the hard part of designing one.
Companion material: the executed notebook stammer.ipynb, and the
radiohead_render scenarios stammer_grid (the dials held still so the
mechanism is audible), stammer_disintegrate (forty seconds of the
performance the object exists for), and stammer_two_seeds.
How it works, in one paragraph
The input is captured continuously into the last few seconds of history. On
a step grid, if the machine is idle, it rolls: with probability density
it grabs the material that just went past, chops it to step divided by
something between 1 and divisions, and plays it back between 1 and
repeats times, each pass with a reverse chance of running backwards. If
jump is open it may reach further back than the bar just played. While a
slice fires you hear the slice; when nothing is firing you hear the input,
untouched.
density and repeats — they are not the same dial
This is the one thing worth internalizing before you patch it. density is
how often the machine grabs; repeats is how long it holds on once it
has. A slice in flight is never interrupted, so repeats is what actually
decides how busy the machine is — and once trains start overlapping, raising
density stops doing anything at all.
Measured off the object's own playing flag, 100 grid points per run. Repeats is the hold.
At density 0.3, going from 1 repeat to 6 takes the machine from 41% busy to 76%. At density 0.9 it is already 90% busy with a single repeat, and 96% with six — the ceiling, where the dial has run out of room.
divisions, reverse, jump — the character
divisions is how finely the grid may be chopped: at 1 you get whole-step
slices, at 8 the machine may cut down to eighths of a step. Because the
divisor is drawn per slice, a high setting gives you a mixture of lengths,
not uniformly short ones — which is what keeps it sounding played rather
than gated.
reverse is drawn per repeat rather than per slice, so a single train can
stagger forwards and back. jump is the reach: at 0 the machine only ever
replays the material immediately past (the classic stutter), and opening it
lets slices come from seconds ago, so the part starts quoting itself out of
order. That is the setting that turns "stuttering" into "disintegrating".
fade is the anti-click — a raised-sine flank on each repeat, exactly zero
at the edges and exactly unity across the plateau. Repeats are sequential
rather than overlapped, so every junction dips to zero. That is deliberate:
the dip is the articulation of a stutter, and a crossfade there would
smear the thing you want to hear.
seed — the dial that is a contract
Every draw — fire, division, repeat count, reach-back, and the per-repeat coin for reverse — comes from a seeded generator in a fixed order. So the same seed and the same moves give the same render, bit for bit. That is not a nicety; it means a take you liked is recoverable, two instances on different seeds decorrelate instead of moving in lockstep, and the tests can assert bitwise equality. A different seed is a genuinely different performance: 89% of samples change.
And at density 0 the dice are never rolled at all — so the seed provably
cannot matter, and the object is a bitwise bypass at any mix. Switched off,
this is not "nearly transparent", it is your input. clear erases the
capture, drops the slice in flight, and rewinds the seeded stream, so the
same seed replays from there.
The material contract
The header says this object wants transient material, and that on a sustained pad a stutter is barely a tremolo. That reads like taste. It is not — it is a property of the material, and it is measurable.
How alike two arbitrary slices of the material are. Re-ordering interchangeable things does nothing.
Every slice of a steady sine looks like every other slice, so shuffling them changes almost nothing you can hear. Slices of a played phrase are all different, so shuffling them is the entire effect. Feed this object drums, plucked or struck strings, consonants — anything whose interest is in when things happen. It re-articulates rhythm that is already in the sound; it cannot invent rhythm that is not.
Recipes
- A grid you can hear:
@step 250 @density 0.55 @divisions 4 @repeats 4 @reverse 0.2. The mechanism, plainly, over a played part. - The disintegration: start at
@density 0.2 @divisions 1 @repeats 1and walk over thirty seconds to@density 0.9 @divisions 8 @repeats 10 @reverse 0.6, tightening@stepfrom 250 to 120 as you go. Then open@jump 1500and the machine starts quoting the wrong bar. - Vocal chop:
@step 125 @density 0.4 @divisions 2 @repeats 3 @fade 6. Consonants are transients; the longer flank keeps it from sounding digital. - Two of them: the same settings on two instances with different seeds, panned apart. They decorrelate by construction — that is what the seed contract buys you.
When it is not the right tool
- Sustained material. See above; it is measured. A tremolo or a gate will do more for a pad.
- Pitched mangling. Slices play at ±1 rate only — there is no pitch
shift and no varispeed here.
tap.shift~transposes;tap.pitchaccum~spirals. - Exact, notated rhythms. The grid is regular but the dice are dice.
tap.808.seq~sequences; this improvises. - Very long repeat trains. A slice reads from the ring, not a private copy, so a train longer than the captured history will start reading fresher material as the write head laps it. Size the object argument to the longest train you intend to fire.
Checkpoint
Capture everything, then on a grid roll dice and re-fire what just went
past. density grabs, repeats holds — and holding is what fills the
timeline. divisions, reverse and jump are the character, and jump is
the one that turns a stutter into a disintegration. The seed is a real
contract: same seed, same performance, bit for bit; at density 0, a bitwise
bypass. And the material contract is measured rather than asserted, which is
the honest way to tell you what to feed it. Every number above lives twice:
as an executed cell in stammer.ipynb and as a pinned scenario in
tests/stammer_test.cpp.
The dirt with two stages
Two objects in this library make things dirty and they are not competing.
tap.overdrive~ is a feedback soft-clipper chasing the Tube Screamer
lineage — the nonlinearity sits inside a loop with a lowpass, so the bass
stays clean and the mids break up first. tap.fuzz~ is the other school:
two clipping stages one after the other, and a tone section that scoops the
middle out. It is the OK Computer-era sound — the dirt on Paranoid Android
and My Iron Lung — and it belongs in this part of the book because, like
the tape echo and the stutter, the interesting settings are the ones you
arrive at by moving something.
The method is not invented here. It is the simplified cascade of Yeh, Abel and Smith's DAFx-07 paper on distortion and overdrive pedals: conditioning filter → memoryless nonlinearity → equalization filter, twice. That paper also supplies the licence for the central shortcut. A real diode limiter is not a static curve at all — it is a lowpass whose pole moves with the voltage across it, and solving that honestly is expensive. Approximating it as a fixed curve between fixed filters is defended there, and measured against real pedals.
What this object is not is a model of a specific pedal. No resistor, capacitor or corner frequency in it is claimed as measured from a unit, and the control names follow the layout that class of pedal conventionally carries rather than asserting what any particular one does.
Companion material: the executed notebook fuzz.ipynb, and the
radiohead_render scenarios fuzz_gain_sweep, fuzz_tone and
fuzz_edge_and_bite.
One curve, two knees
Both stages share a single clipping family — tanh(kx)/tanh(k) — normalized
so that full scale in is full scale out at every knee. That normalization
is what lets the knee be a character control instead of a hidden volume
control.
The knee sharpens the corner without moving the ceiling.
The first stage takes a soft knee and most of the gain (the op-amp-ish
stage); the second takes a harder one at unity (the shunt limiter). edge
sweeps the second stage's knee from a gentle limiter toward something close
to a hard corner.
gain — and why the floor is below unity
The knob sweeps the first stage's drive. Its floor sits below unity deliberately, and the reason is the most useful thing in this chapter if you ever build a cascade of your own.
The tanh family's small-signal gain is k/tanh(k) — greater than one, and
growing with the knee. Put a fixed ×2.2 in front of a knee-3 curve and the
second stage sees an effective ×6.6, which means it is fully clipped before
the gain knob leaves zero. That is exactly what the first version of this
kernel did. It sounded like a distortion at every setting, which is precisely
why listening did not catch it and a measurement did.
Left: the gain knob after retuning — harmonic content sweeps 0.010 to 0.358. Right: asymmetry is what makes even harmonics.
asymmetry — the even harmonics
A symmetric curve is an odd function, so it can only make odd harmonics. DAFx-07 points out that a real op-amp stage clips lopsided, and that this is where a pedal's even-order content comes from — which is the whole reason this control exists. Turn it up and the even/odd ratio climbs from essentially zero to about 0.55.
It costs no DC. The bias is applied inside the curve and corrected at the stage output, so however lopsided the setting, silence in is exactly silence out — no pedestal, no thump when you stop playing.
bass, treble, contrast — the voicing
Three linear filters entirely outside the nonlinearity: a low shelf, a high
shelf, and a mid scoop whose depth is contrast. On this class of pedal the
voicing section is most of the identity — the scoop is the sound people mean
when they describe it — so it is a first-class part of the object rather
than an afterthought bolted on at the end.
oversample — and a default that was wrong twice
A static curve makes harmonics without limit, so anything above Nyquist folds back. The clipper pair therefore runs oversampled. Everything about this control has been re-measured, because the first two conclusions drawn from it were wrong, and wrong in the same way.
First, the anti-alias filter here is 8th order, where the rest of the house uses 4th. Measured in this kernel the 4th-order pair is not steep enough — alias energy at 4× came out worse than at 2×.
Second, the chain is a cascade of 2× stages — one doubling, one filter, repeated — rather than a single zero-stuff by the whole factor. That is what finally made more oversampling mean less aliasing. The single-stage chain left N−1 images for one filter to suppress at a corner that got tighter with every doubling, and the residue intermodulated in the clipper into exactly the non-harmonic junk the probe measures. Cascading removes the reversal outright, and where 4× and 8× used to be merely adequate they are now two to four orders of magnitude cleaner. It costs about 5 % more CPU at 8×.
Third — and this is the part worth taking away — the old default came from a single test tone. Every number in the original write-up was measured at 3733 Hz, and 2× happens to look best there. Swept across tones, 2× collapses above about 6 kHz; at 10.5 kHz it is worse than not oversampling at all, because the clipper's low harmonics already exceed the base Nyquist and one doubling does not move them out of the way.
| input tone | 1× | 2× | 4× | 8× |
|---|---|---|---|---|
| 3733 Hz | 1.2e-1 | 3.0e-5 | 2.1e-5 | 2.2e-5 |
| 5171 Hz | 1.5e-1 | 3.3e-4 | 2.0e-7 | 1.9e-7 |
| 6421 Hz | 8.5e-2 | 3.2e-2 | 3.9e-7 | 3.6e-7 |
| 8123 Hz | 9.0e-2 | 7.8e-2 | 1.1e-3 | 2.0e-5 |
| 10499 Hz | 1.5e-1 | 1.7e-1 | 1.2e-5 | 1.6e-6 |
So: 4× is the default. More is never worse now, and 4× is indistinguishable from 8× below about 7.5 kHz. Above that, harmonics start folding inside the 4× band before decimation — 8123 Hz in the table is that happening — and 8× is worth the extra 1.4 % of a core.
Use 2× only if you have measured your own material and it holds up there. It is kept because it is cheap and because on a bass-heavy source it is fine, not because it is good.
Recipes
- Edge of breakup:
@gain 0.3 @edge 0.2 @contrast 0. @bass 0.. Barely dirty; a boost with attitude. - The scoop:
@gain 0.8 @edge 0.6 @contrast 1. @bass 0.4 @treble 0.2. The sound the control is named for. - Lopsided and mean:
@gain 0.9 @edge 1. @asymmetry 0.7 @oversample 8. Hard knee plus even harmonics, and 8× because a hard knee on a bright source is exactly where the top octave folds. - Into the echo:
tap.fuzz~→tap.tapecho~with the echo's@drivelow. Two saturators in series get muddy fast; let the pedal be the dirt and the tape be the space.
When it is not the right tool
- Amp-like breakup.
tap.overdrive~keeps the bass clean by design; this object does not, and hard settings will get woolly on a bass-heavy source. - Subtle warmth. Two stages is a lot of stages. At low gain this is a
clean boost with a tone stack, which is fine, but
tap.overdrive~is the better instrument for gentle. - A specific pedal. This is that pedal's class. If you need a named unit, this is not it and does not pretend to be.
Checkpoint
One clipping family with a knee control, cascaded twice, into a voicing
section that scoops the middle. The gain knob's floor is below unity because
small-signal gain compounds through a cascade — a lesson that cost this
kernel one wrong first draft. asymmetry is the even-harmonic control and
costs no DC. And the oversample setting is a
measurement twice corrected: cascaded 2× stages, because a single zero-stuff
by N was what made bigger measure worse — and a default of 4× rather than 2×,
because the old default had been generalized from one test tone. Every number here
lives twice, as a cell in fuzz.ipynb and as a pinned scenario in
tests/fuzz_test.cpp.
Two hands on the same tape
tap.stammer~ and tap.scrub~ record the same way. Both keep a rolling
tape of what just went past — the same capture, literally the same code,
not a second copy of it — and both put a read pattern on top of it. The
stutter's pattern is a slicer with dice. The scrub's is a pad you drag.
That is the whole difference, and it is the difference between a machine that decides and a machine you play. The stammer is a die you load; the scrub is a surface you push around, in the Kaoss-pad school of instruments where an XY surface over live capture is the entire interface. It belongs in this part of the book for the same reason the tape echo does: nobody sets this object up and walks away from it.
It is an original design in the granular / brassage tradition (Roads, Microsound, MIT Press 2001) — not a port, and not a reconstruction of any product. No preset, timing or parameter value in it came from a piece of hardware.
Companion material: the executed notebook scrub.ipynb, the pinned
scenarios in tests/scrub_test.cpp, and the radiohead_render scenes
scrub_gesture and scrub_freeze.
The two axes are actually two axes
position is how far back the playhead sits, as a lag in milliseconds
behind the live edge. pitch is transposition in semitones. On tape those
would be the same knob — moving the head is the pitch change — and the
whole point of doing this with grains is that here they are not. Hold the
position and sweep the pitch and the material transposes without going
anywhere. Sweep the position at a fixed pitch and you rake through the last
few seconds without the tape rising or falling.
drift is the third one, and it is the playhead's own motion through the
tape in playback-rate units: 1. runs forward at the speed the recorder is
writing, 0. holds station, negative runs backwards. Set @drift 1. and
let go of the position and the scrub is a delay; set @drift 0. and it is a
freeze that you can still transpose.
The identity underneath it
Grains are Hann-windowed and fired every size / overlap samples. Hann
overlap-adds to exactly 1 at that hop, so with the pitch at unity, spray
at zero and the position held on a whole sample, the scrub is the input,
delayed, to floating point — 4.4e-16 in the pinned test.
Left: held still at unity, the object is a delay and nothing else. Right: the window sum that makes it one.
This matters more than it sounds. Everything else the object does is a departure from a plain delay, and a departure is only trustworthy if you know the thing it departs from is exact. When the position drags, when the pitch moves, when spray scatters the origins — those are the object working. If the still case were approximate, you could not tell them apart from noise.
overlap 1 leaves gaps between grains, which is the dipping curve in that
figure. That is a chopped, gated texture rather than a defect, and it is
worth having; it is just not the setting the null lives at.
What transposing costs, honestly
Reading tape at a rate the write head does not share means the read pointer drifts away from where the position says it is, and it has to be pulled back or the position stops meaning anything. Every pull-back is a splice between two grains reading material a little apart.
What that costs is not the pitch. Swept over seven fundamentals and seven intervals, 98.8 % of a perfect shifter's energy lands within ±15 Hz of the transposed pitch — worst case 91.7 %. The note is where you asked for it.
What it costs is concentration. The band holds a narrow comb rather than one clean line: 92.0 % as concentrated as a clean shift, 75.0 % at its worst.
The pitch goes where you put it. What the splices take is focus.
Audibly that is a warble, and it is the classic single-delay-line
pitch-shifting artifact rather than anything peculiar to this kernel. If you
want the warble gone, spray trades the comb for a broadband smear, which
some material prefers. If you want a clean shift, this is the wrong object —
see below.
freeze, and what it does not stop
freeze stops the recorder. The playhead keeps going, so the position now
addresses fixed tape and the grains loop the same window: a granular hold
you can still scrub, transpose and drift through. It does not stop time
inside a grain — a grain in flight when freeze engages was already
scheduled, and it finishes.
spray and seed
spray scatters each grain's origin randomly back from the position. At
exactly 0 the dice are never rolled, so the seed provably cannot matter —
the same contract tap.garden~ and tap.stammer~ carry, pinned by the same
kind of test. With spray up, the same seed and the same moves give the same
render bit for bit, and two instances decorrelate by seed alone.
Recipes
- The pad:
@drift 0. @size 80 @overlap 2 @mix 100, then ridepositionwith a signal. The default instrument. - Granular freeze:
@freeze 1 @drift 0. @size 120 @spray 40. Hold, then movepitchfor a chord that was never played. - Backwards tape:
@drift -1. @pitch 0.— the playhead walking against the recorder. - Chopped:
@overlap 1 @size 40. Gaps between grains, on purpose. - Into the diffuseur:
tap.scrub~→tap.palme~with the palme's@mixaround 40. The strings sustain what the scrub chops.
When it is not the right tool
- Clean transposition. The warble above is inherent to the method.
tap.shift~andtap.pitchaccum~are the objects built for that job — with one caveat worth stating plainly: measured on this same sweep,tap.pitchaccum~retained mean 0.907 of band energy against the scrub's 0.988, worst 0.633 against 0.917. Its two-tap crossfade also puts its strongest spectral line a few hertz beside the intended pitch, which is filed as issue #33. None of that makes it the wrong object — it is a shimmer, and shimmer is what those sidebands are — but audition before you assume it is the transparent one. - A tidy delay.
tap.delay~andtap.tapecho~cost far less and do not window anything. - Slicing to a grid. That is
tap.stammer~, on the same tape.
Checkpoint
One tape shared with the stutter, one grain scheduler on top of it, and two
axes that stay independent because grains let them. The still case is a
bit-exact delay, which is what makes every departure from it legible.
Transposing warbles, and the warble is measured rather than apologized for:
the note holds to 98.8 % of a clean shifter's energy, and 92.0 % of its
focus. Every number here lives twice, as a cell in scrub.ipynb and as a
pinned scenario in tests/scrub_test.cpp.
Loudspeakers you can play
The Ondes Martenot does not have a loudspeaker. It has a rack of them, and the player chooses. Beyond the plain cabinet — the principal — Maurice Martenot built resonating diffuseurs whose entire job is to colour the signal with a physical body: the métallique (1944–45, patented 1947), a gong driven by a motor transducer, and the palme (1949–50), an electromagnet driving twelve metal strings stretched on a soundboard.
tap.metallique~ and tap.palme~ are those two, and they ship as
standalone effects rather than as something hidden inside tap.ondes~,
because the interesting thing about a resonating loudspeaker is that it does
not care what you put through it. A guitar into the palme is not what
Martenot had in mind and it is the best reason to have the object.
Najnudel, Hélie, Roze and Boutin (IEEE/ACM TASLP 28, 2020) name the diffuseur as the stage that "converts the electrical waveform into sound and in turn modifies its spectral content". Wijnand, Boutin, Jossic and Maniguet (Forum Acusticum 2023) describe the instruments and measure the transducer. Everything below traces to one of those two, or is labelled as a recreation.
Companion material: the executed notebook diffuseur.ipynb,
tests/diffuseur_test.cpp, and the radiohead_render scenes
metallique_stages and palme_halo.
Driven, not struck
tap.chime~ and tap.garden~ already carry this library's modal machinery
— mode ratios, doublet splitting, per-mode decay — and it carries over here
intact. What does not carry over is the strike. There is no trigger in
either of these objects and no decay envelope. A diffuseur is excited
continuously by whatever is going through it and rings at its own rates,
which is tap.5comb~'s sustained-resonance situation rather than the
chime's.
Practically, that is the difference between an object you fire and an object you feed.
The order is the argument
The electrical signal reaches the transducer first, and the transducer's motion is what excites the body. So the nonlinearity sits upstream of the resonator. Drive the transducer hard and you are pushing a distorted waveform into a gong — which is a different sound from distorting a gong.
That claim is pinned rather than asserted: a null test in the kernel checks
that a whole cabinet is bitwise identical to transducer → body wired by
hand, and that the reverse wiring differs by 28 % of peak. It is not a
subtlety you have to take on faith, and it is not a subtlety you can hear
your way past.
The métallique
Eight modes at the free circular plate's transverse ratios — Rayleigh's
classical Chladni set at Poisson 0.3, 1 : 1.730 : 2.328 : 3.910 : 4.110 : 6.300 : 6.710 : 7.340 — each split into a slowly beating doublet.
Left: where the modes are. Right: the body answering a sweep, which is how you actually meet it.
pitch places the lowest mode and the rest follow. decay is the
fundamental's T60 — long is a drone, short is a plate reverb. tilt decides
how much faster the upper modes die than the fundamental, and brightness
weights them. The weights sum to exactly 1 and each mode has unit peak gain,
which is why there is no limiter on the output and no DC blocker either: the
body is bounded by its input, by construction.
The palme
Twelve strings, each a damped delay loop, on one board.
Twelve, not twenty-four. Widely copied build pages say two banks of twelve; the peer-reviewed source says twelve, and this object follows the peer-reviewed source.
Their tuning is not published anywhere found, so it is a control: @tuning 0 lays them out chromatically across an octave from root — a string for
every pitch class, so the board answers whatever you play — and @tuning 1
puts the harmonic series on the root, which is a drone that answers one key.
Feed it a tone, take the tone away, measure what is left. Every one of the twelve strings rings at least 4.4× harder at its own pitch than between them.
damping is how fast a string loses its upper partials — low values are
felt cloth on the strings. detune scatters the strings against each other
by a fixed, deterministic amount in cents, because no two strings on a real
board are in perfect relation.
The transducer
Wijnand et al.'s point about the early diffuseurs is that they use a moving-iron driver whose operating principle is inherently nonlinear — Thiele–Small does not describe it — so a diffuseur modelled as a pure resonator is missing a documented stage.
What is modelled here is that principle, not a fit to a measurement. In a
moving-iron motor the force follows the square of the gap flux, so with a
bias current I₀ and signal i the force carries a term in (I₀ + i)²
whose residual i² makes second-harmonic distortion that grows with drive.
That is asymmetry: the transducer's own even-harmonic signature, and the
only part of these objects that is nonlinear by citation.
saturation is the honest exception. A squared law is expansive and
something has to bound it, so there is a soft clipper after it — a
modelling necessity, not a measured stage, and its coefficient is a knob
rather than a number from a paper. At 0 it is exactly linear.
What these are, and are not
The instruments, their dates, their excitation and their transducer type are peer-reviewed. The modal data is not. No ondes-specific measurement of either body exists in any source obtained, so the plate comes from Fletcher & Rossing's free circular plate and the strings from the harmonic series. Both bodies are therefore recreations of the general physics, not models of Martenot's instruments. Nothing here was fitted to a recording, a measurement, or a photograph.
There is also no radiation model — no directivity, no cabinet, no soundboard
resonance of its own. The output is the body's modal response, not a room.
And the strings are ideal: a real steel string is stiff and its partials
stretch sharp, and that dispersion is not modelled. detune scatters
strings against each other, which is a different thing and does not stand in
for it.
Recipes
- A guitar into the palme:
tap.palme~ @root 110 @tuning 0 @decay 8 @mix 45. The halo underneath everything you play. The reason these ship standalone. - The instrument, assembled:
tap.ondes~→tap.palme~ @mix 60. What Martenot actually had. - Gong reverb:
tap.metallique~ @pitch 180 @decay 1.5 @tilt 1.2 @mix 35. Short decay turns the body into a plate. - A drone you drive:
tap.metallique~ @decay 20 @drive 3 @asymmetry 0.5 @saturation 0.4 @mix 100. Hard into the transducer, which is upstream, so it is a distorted waveform ringing a gong rather than a distorted gong. - The one to be careful with:
tap.palme~ @level— twelve resonant loops add up, and a driven board can be much louder than what went into it.
When it is not the right tool
- A reverb. These are twelve strings and eight modes. They are pitched, and they will impose their pitches on anything you send.
- A model of Martenot's own diffuseurs. See above: this is the physics of the general case, and the difference is stated rather than glossed.
- Clean sustain.
tap.5comb~is the sustained-resonance object without a nonlinear driver in front of it.
Checkpoint
Two loudspeakers with bodies, shipped as effects because a resonating cabinet does not care what drives it. Driven rather than struck, so no trigger and no envelope. The transducer is upstream of the body and a bitwise null test pins that it is — 28 % of peak says the order is audible. Every mode has unit peak gain and the weights sum to 1, so the body needs no limiter. And the bodies are recreations of published physics rather than measurements of Martenot's instruments, which is a limitation stated here and in the header rather than left to be discovered.
The instrument that is not a synthesizer
Three objects in this chapter — tap.ondes~, tap.triode~ and
tap.touche~ — and one instrument. The Ondes Martenot, 1928, the thing
Messiaen wrote for and Jonny Greenwood plays: a keyboard you can also play
with a ribbon on a ring, and a pressure key in the left hand that is the
whole dynamic range of the instrument.
The plan for this family assumed it would be an oscillator with waveform switches. It is nothing of the kind, and finding that out changed every decision below.
The Ondes Martenot is heterodyne. Two oscillators run near 80 kHz, one fixed and one moved by the ribbon; they are summed, and the note you hear is the envelope of their beating. Najnudel, Hélie, Roze and Boutin, who modelled instrument No. 169 stage by stage (IEEE/ACM TASLP 28, 2651–2660, 2020), measure those oscillators at about 0.03 % second harmonic even coupled to the rest of the circuit. They are essentially pure sinewaves. Every bit of the instrument's character therefore comes from what happens after them: the demodulator, two valve stages, the intensity key, and the diffuseur.
Companion material: the executed notebooks ondes.ipynb and touche.ipynb,
tests/ondes_test.cpp and tests/touche_test.cpp, and the
radiohead_render scenes ondes_stages, ondes_ribbon, ondes_diffuseurs
and triode_tubes.
The biggest source of harmonics is not a valve
Two oscillators of equal amplitude sum to an envelope of 2|cos|. That is
not a sinusoid. Its Fourier series puts the second harmonic 14.0 dB
below the fundamental, the third 21.3 dB down and the fourth 26.4 dB
down — a substantial harmonic series generated before anything nonlinear
touches the signal.
The demodulator is the instrument's largest single source of harmonics, and it is upstream of every valve.
This is why tap.ondes~ does not synthesize a difference tone. Generating
the note as a sinewave and distorting it afterwards would throw away the
part of the timbre that arrives for free — and it is an easy mistake to
make, because the circuit paper does say the oscillators can be replaced
by a sinewave generator. That licence applies to the oscillators, not to
the demodulator.
The carrier is not simulated either, and that is not a compromise. For
amplitudes 1 and depth the envelope is exactly sqrt(1 + depth² + 2·depth·cos φ), so the 80 kHz disappears from the arithmetic rather than
being approximated away. Running the published RC detector on that closed
form matches a full heterodyne-plus-diode-plus-RC simulation to within
0.10 dB on every harmonic at every pitch tried.
depth is that second amplitude, and it turns out to be the cheapest real
timbre control in the object. At 1 the envelope closes completely and the
series is full; below 1 it never closes and the tone thins toward a
sinusoid. It is a mismatch between two real oscillators, not an invented
knob.
detect — and why the instrument thins as it climbs
The detector is the published one: a triode grid near zero bias conducts on
positive half-cycles and charges instantly, and R4 × C21 = 1 MΩ × 200 pF
discharges it — a 200 µs time constant, which is @detect 0.2.
That single number carries the instrument's pitch character, because an RC that slow cannot follow a fast envelope back down. Measured here, the second harmonic runs from −14.0 dB at A2 to −19.3 dB at A6, and the level falls 2.0 dB across those five octaves. The ondes gets purer and quieter as it goes up, and it does so for a reason you can point at in a schematic.
The ribbon is linear in semitones
The circuit paper's Eq. 7 gives the variable oscillator's capacitance
against ribbon displacement, and what falls out is
f = 55 Hz · 2^(d / 12·d₀).
So @ribbon is semitones above A1, not Hz. A hand moving at constant
speed makes a constant-rate glissando; nothing quantizes, and nothing
should. This is why an ondes glide sounds the way it does, and it is the one
place where taking the units from the paper rather than from convention
changes how the object feels to play.
tap.touche~ — 50 dB in four and a half millimetres
The intensity key is a graphite-and-mica powder bag working as a rheostat: compress it and the number of conducting bead paths rises, so resistance falls. Messiaen called it the instrument's greatest invention. What the player feels is a well-chosen nonlinear spring.
The curve in this object is not modelled and not fitted. Quartier, Meurisse, Colmars, Frelat and Vaiedelich (Acta Acustica 101(2), 421–428, 2015) measured finger force, key displacement and sound simultaneously on instrument No. 320, and published the boundaries of the six musical nuances across the key's travel. Those seven points are the object, interpolated with monotone cubic segments that pass through every one of them.
Seven measured points, and the shape between them. The straight line is what a fit would have thrown away.
Three things follow, and each is a decision the paper made rather than this object:
- Position, not force and not velocity. The paper states explicitly that the intensity depends on displacement, and not on the speed of the gesture. A static memoryless map is the finding, not a simplification.
- 50 dB over about 4.5 mm, from 4.3 mm (the instrument's noise floor) to 8.8 mm. The paper notes most traditional instruments rarely exceed 25 dB of per-note dynamic range.
- The shape is not a line. Equal 8.3 dB steps take displacement steps of 1.0, 0.6, 0.5, 0.4, 0.5 and 1.5 mm. It steepens through the middle and flattens hard at the top.
And the thing that surprises everyone who patches it: on a 0–1 control, roughly the bottom 45 % of the travel is silent. That is not a dead zone in the object. It is the key's own first phase — the elastic strip bending before it reaches the powder bag — and it is exactly why the instrument can be attacked so sharply, because the useful 50 dB lives in the 4.5 mm right after it.
tap.triode~ — the stage is a citation
The valves are where the rest of the character is, and there was nothing to invent. The circuit paper does not merely mention a tube model: it names the enhanced Norman Koren model (Koren, Glass Audio 8(5), 1996, with Cohen & Hélie's grid-current extension, AES 129, 2010), writes out its equations, and publishes parameter sets fitted to the actual valves in ondes No. 169 in its Table II — 6F5 in the oscillators, 6C5 in the demodulator and preamplifier, 2A3 in the power amplifier — along with each stage's supply voltage, cathode resistor and plate load.
A stage is then the static solution of the load line, which is a memoryless
nonlinearity in exactly the DAFx-07 sense tap.fuzz~ uses. Where the fuzz
reaches for a tanh, this one solves a valve.
Left: the published operating point, solved. Right: the stages invert, and they are visibly lopsided.
Two properties matter before you patch tap.triode~ on its own:
- It inverts, as a real common-cathode stage does. That is not cosmetic. The valve's asymmetry acts on whichever side of the waveform reaches its grid, so the sign decides which half gets bent.
- It is strongly asymmetric. At the demodulator's operating point, equal grid swings either way give plate swings in a 2.17 : 1 ratio. That ratio is where a triode's even harmonics come from.
drive is normalized out of the level — the gain-staging lesson tap.fuzz~
learned the hard way, applied here from the start — so turning it up gets
dirtier rather than louder.
drive on the voice, and where it starts from
The valves add to a signal that was already rich. The floor is the demodulator's.
The important thing in that figure is the dotted line. At @drive 0. the
tone still measures 0.221 of harmonic content, because the demodulator made
it. The knob sweeps 0.221 → 0.344, monotonically, without the level running
away.
The two controls that are choices
Most of this object is a citation. Two controls are not, and both are labelled as such because both measure as audible.
keyplacement— the paper's five stages do not include the intensity key, so where it sits is undetermined. After the valves (the default) it is a clean output law: pressure is level. Before them, pressure drives the valves: soft is clean and hard is dirty. The two differ by about 0.09 of total harmonic content at a half-press.polarity— the two valve stages are coupled through a transformer whose winding sense is not in the source, and the sign decides which side of the waveform the preamplifier's asymmetry acts on. Worth about 0.12.
power, and taking the authors at their word
The 2A3 power stage is off by default, following the paper: they measure almost 5 % second harmonic there, but report its contribution as much less important than the two stages before it, and drop it for real-time.
Measured here, switching it on moves total harmonic content from 0.248 to 0.251 and the second harmonic by 0.1 dB. They were right, which is why it is a switch rather than a deletion.
oversample
The nonlinear chain runs oversampled. Worst non-harmonic energy relative to the fundamental, at 1× / 2× / 4× / 8×:
| tone | 1× | 2× | 4× | 8× |
|---|---|---|---|---|
| 587 Hz | −79.3 | −91.2 | −104.5 | −103.8 |
| 1175 Hz | −65.8 | −77.2 | −90.6 | −92.5 |
| 1760 Hz | −57.6 | −70.9 | −81.1 | −82.2 |
| 2637 Hz | −51.1 | −61.4 | −71.8 | −83.8 |
| 3520 Hz | −45.4 | −56.8 | −67.0 | −74.2 |
Every doubling is worth about 12 dB up to 4×; past that it is worth 7–12 dB at the top of the range and nothing at the bottom, where the measurement has already bottomed out. Never worse. 4× is the default because that is where the cost stops buying uniformly; 8× is there for anyone playing the top octave hard.
Readers of the tap.fuzz~ chapter will notice this used to be the opposite
of what that object measured. That was not a contradiction — it was the clue
that fixed the fuzz. The appendix explains how.
What is missing, deliberately
The real instrument has waveform registers — switchable timbres. Their filter shapes are in none of the sources obtained, and inventing them is the one thing this object will not do.
There is also no diffuseur in tap.ondes~, because that is
tap.metallique~ and tap.palme~, and patching one after the other is how
the instrument works anyway.
Recipes
- The instrument:
tap.ondes~→tap.palme~ @mix 60. Ribbon and key on signals; that is the whole performance surface. - Ribbon on a slider:
@ribbontakes a signal, and aline~from 0 to 36 over four seconds is a three-octave glissando that sounds like one because the law is linear in semitones. - The key alone:
tap.touche~on any source. It is a published expressive gain law, and nothing about it is ondes-specific once it is detached. - Thin and pure:
@depth 0.4 @detect 0.6 @drive 0.. The envelope never closes and the detector smooths what is left. - Dirty on hard presses:
@keyplacement 1 @drive 4 @polarity -1. Pressure drives the valves. - A valve on a guitar:
tap.triode~ @tube 2 @stage 2 @drive 6— the 2A3 power stage, used for something it was never in this instrument for.
When it is not the right tool
- A subtractive synth. There is no filter, no envelope generator and no waveform selection here. It is one voice with a ribbon and a key.
- A polyphonic anything. The instrument is monophonic; so is this.
- A specific recording. The valve parameters are a fit to one instrument's tubes, and tube-to-tube spread in 1930s valves is wide.
Checkpoint
A heterodyne instrument whose oscillators are nearly pure, so the character
lives downstream: a demodulator whose 2|cos| envelope makes more harmonics
than either valve does, two valve stages that are a published model with
published parameters, and a pressure key that is a published measurement
interpolated rather than fitted. The ribbon is linear in semitones because
Eq. 7 says so. Two controls are choices rather than reconstructions and are
labelled as choices. The waveform registers are missing on purpose. Every
number here lives twice, as a cell in ondes.ipynb or touche.ipynb and as
a pinned scenario in tests/ondes_test.cpp or tests/touche_test.cpp.
Making the machine talk
The vocoder is audio's oldest identity theft: take the shape of one sound
and wear it over the body of another. Speech works because your mouth
sculpts a moving spectral envelope; a vocoder measures that envelope on one
signal (the modulator — usually a voice) and stamps it onto another (the
carrier — usually a synth), and the synth talks. tap.vocoder~ is the
classic architecture: a 24-band channel vocoder, time-domain, no FFT. This
chapter is how to wire it and — mostly — how to choose the two signals, which
is nine tenths of vocoding.
Companion material: the reference page and help patcher in the TapTools-Max package; the kernel's Catch suite pins the structural behavior quoted below.
The machine, in one pass
Two identical banks of 24 bandpass filters, log-spaced from 50 Hz to 12 kHz (RBJ constant-peak biquads — unconditionally stable across the range). The modulator goes through one bank; a per-band envelope follower measures each band's level. The carrier goes through the other bank; each carrier band is multiplied by the matching modulator envelope; the bands are summed. That's the whole machine — which is why its behavior is so predictable:
- A silent carrier is silence, no matter what the modulator does (pinned by test): the modulator only ever gates; every sample you hear is carrier.
- Gain is exactly linear (pinned): the vocoder adds no nonlinearity of its own.
- A silent modulator decays to silence at the follower rate — the vocoder "lets go" of the carrier the way the voice lets go of a word.
The wiring
Modulator in the left inlet, carrier in the right. Getting these backwards is the classic first-patch bug, and it sounds like it: a synth "speaking" your voice is right; your voice weakly filtered by a synth is backwards.
Two identical banks meeting at 24 multipliers. Envelopes gate the carrier; modulator audio never reaches the output.
The knobs, one by one
q — intelligibility vs. smoothness
The bandwidth of all 48 filters. Narrow (high q) separates the bands cleanly — crisper consonant detail, more "robot" — but thins the carrier between band centers. Wide (low q) overlaps the bands into a smoother, duller blend. The classic hardware vocoders sat toward smooth; intelligibility came from performance, not q.
response_interval — how fast the mouth moves
The envelope followers' period in ms. Short tracks every consonant — crisp, maximally intelligible, and a little nervous. Long smears syllables into pads — the "choir" setting. This knob is the vocoder's attack and release; 20–50 ms speaks, 200+ ms sings.
gain
Makeup level, since a band-multiplied signal usually lands quieter than either input. Linear, boring, necessary.
sibilance — the built-in s and t budget
The classic channel-vocoder unvoiced path (Dudley's lineage): a seeded
internal noise source blended into the carrier of the bands above
~4 kHz, still gated by the modulator's envelopes — so consonants articulate
even over a dull carrier, and only when the modulator actually has
high-band energy (pinned: a silent carrier with an HF-rich modulator
speaks at sibilance 1; a low-only modulator stays quiet). At the default
0 the original silent-carrier contract holds exactly, bit-identical —
turning it up deliberately relaxes that contract for the top bands. The
noise is deterministic per seed, family doctrine.
mix — the synth under its own robot voice
Equal-power blend of the dry carrier against the vocoded output — the classic parallel move (the pad fades in under itself talking). Endpoints are exact: 100 is bit-identical wet, 0 returns the carrier untouched.
Choosing the two signals (the actual craft)
- The carrier must have energy where the modulator has bands. The eternal
vocoder failure is a dull carrier: a mellow sine pad gives the high bands
nothing to gate, and consonants vanish. The house answer is upstairs in
this book — a
tap.vco~saw stack (harmonics forever, and the analog section keeps it moving) is a nearly ideal carrier; noise (tap.noise~) blended in restores the s and t sounds that even a saw can't carry. - The modulator wants articulation, not fidelity. Overdriven, compressed, even cheap-microphone speech vocodes better — what matters is envelope contrast between bands, not beauty.
- Nobody said voice. Drums modulating a pad turns the pad into rhythm; a cello modulating noise is a ghost. The machine imposes any moving envelope on any body.
Recipes
- The talking synth: speech → left;
tap.vco~saw stack + 10 % noise → right;response_interval 30,qmiddling, and enunciate like you're annoyed. - The choir: sustained "aah"s → left; detuned saws → right;
response_interval 250. Consonants don't matter; vowels are the chord. - Rhythm transfer: a drum loop → left; anything sustained → right; short
response_interval. The drums play the pad.
When it is not the right tool
- Pitch correction or transposition — a channel vocoder never changes the
carrier's pitch; it only shades its bands. Pitch is
tap.shift~/tap.pitchaccum~territory. - High-fidelity cross-synthesis. Twenty-four bands is a voice, not a
spectrograph; for surgical spectral morphing you want FFT-domain tools
(
tap.spectra~is the start of that corridor). - Formant preservation while shifting — related, but a different machine:
tap.harmony~, which multiplies the voice itself instead of wearing it over a carrier.
Checkpoint
Two matched 24-band banks, 50 Hz–12 kHz: the modulator's per-band envelopes
gate the carrier's bands, and everything you hear is carrier. q trades
crispness against smoothness, response_interval is the mouth's speed, and
the craft is almost entirely in feeding it a bright, busy carrier and an
articulate modulator. The machine is simple; the casting is everything.
A gate for every bin
A noise gate is a bouncer with one rule: too quiet, you don't get in. Useful,
but blunt — when the signal plays, all the noise under it walks in too, and
when the signal stops, the gate slams on room tone. tap.nr~ hires a
thousand bouncers instead: it transforms each STFT frame and applies the
threshold per frequency bin, so the quiet bins between your signal's
partials close while the loud ones stay open. Hiss disappears from the gaps
in the spectrum, not just the gaps in time. This chapter is the two knobs,
the two costs, and the one artifact to listen for.
Companion material: the reference page and help patcher in the TapTools-Max package; the kernel's Catch suite pins the reconstruction claims below.
The contract: transparent until it isn't
The object runs its own STFT — Hann window, 4× overlap, COLA-normalized
overlap-add — and the engineering contract is pinned by test: with the gate
open, the output reconstructs the input exactly (below 10⁻⁶), delayed by
one FFT frame. Whatever tap.nr~ does to your sound, it is doing it on
purpose with threshold and slope; the machinery itself is transparent.
Also pinned: a tone below threshold is strongly attenuated; a tone above
passes untouched.
The pump this object runs on — tap.nr~ is this scaffold with a per-bin downward expander in the middle.
The knobs, one by one
threshold — where quiet begins
The per-bin level (linear amplitude) below which a bin is attenuated. The
craft: set it between your noise floor and your signal's quietest partials.
Play the noisy source silent for a moment, raise threshold until the noise
just vanishes, then stop — every further dB starts eating signal.
slope — how hard the door closes
The soft knee. 0 passes everything (bypass by another name); low values fade
bins gently as they approach the threshold; high values approach a hard
per-bin gate. And here lives the genre's famous artifact: push slope hard
with threshold high and bins near the boundary flicker open and shut frame
by frame — musical noise, a watery, birds-in-the-pipes chirping. The cure
is almost always a gentler slope and a lower threshold, accepting a little
noise instead of a lot of artifact. Half the craft of spectral gating is
knowing when to stop.
FFT size — resolution vs. smearing (and the latency)
The frame size trades three things at once:
- Frequency resolution: bigger frames separate closely spaced partials from noise between them — better gating for dense, tonal material.
- Time smearing: bigger frames blur transients; a gate decision spreads across the whole frame. Percussive material wants smaller frames.
- Latency: exactly one FFT frame, by construction. 2048 samples at 48 kHz is 43 ms — fine on a mix bus, noticeable on a live input.
Recipes
- Location dialog cleanup: moderate frame,
thresholdfound by the silent-passage method above,slopeas low as removes the hiss. Listen to the pauses — that's where both the win and the artifact live. - Synth-line de-hiss: tonal material with stable partials is the best case — bigger frames, and the gate closes every bin the notes don't own.
- Creative abuse: absurd
thresholdwith a hardslopeisn't repair, it's an effect — the signal reduced to its loudest spectral bones. The artifact becomes the instrument.
When it is not the right tool
- Noise under the signal, not beside it. A gate — even per-bin — only removes noise where the signal isn't. Broadband hiss sharing bins with a broadband source needs subtraction/statistical methods, a different machine.
- Hum and buzz. A 50/60 Hz family is a few known frequencies; surgical
notches (
tap.filter~) beat a thousand bouncers who all have to guess. - Time-domain gating with musical envelope shaping — attack/hold/release on the whole signal is a classic gate's job, and it doesn't smear transients.
Checkpoint
An STFT expander: per-bin thresholds close the spectrum's quiet gaps, the
machinery reconstructs bit-faithfully when open (pinned below 10⁻⁶), and the
price is one frame of latency plus the musical-noise artifact that appears
exactly when threshold and slope are pushed past honest. Find the floor,
close the door gently, and stop while the pauses still sound like air instead
of water.
The spectrum, re-plumbed
Every process so far in this book treats the spectrum with respect: filters
shade it, gates prune it, vocoders dress it up. tap.spectra~ re-plumbs it.
Each output bin k is filled from input bin round(k · remap) — the
spectrum's contents redistributed by a rule with no acoustic justification
whatsoever. It is the one object in this book whose purpose is to sound like
nothing in nature, and it is honest about it: the reference page has called
it an "ultra-non-linear effect" since 2002. This chapter is what the rule
does, why the results are inharmonic almost everywhere, and how to drive an
effect whose sweet spots are narrow and strange.
Companion material: the reference page and help patcher in the TapTools-Max package; the kernel's Catch suite pins the two anchor behaviors below.
The rule, and its two pinned anchors
Inside the object's own STFT (the same Hann/4×-overlap engine as tap.nr~),
the lower half of the output spectrum is assembled by reading input bins at
k · remap, and the upper half is mirrored to keep the spectrum Hermitian —
so the output is always real, whatever violence the remap did. Two behaviors
are pinned by test:
remap 1is the identity: the output reconstructs the input exactly, delayed by one FFT frame. Transparent machinery, like its sibling.remap 2moves input bin 2k to output bin k — the spectrum compressed toward the bottom: content from twice the frequency lands at half.
Why almost everything comes out inharmonic
Pitch shifting scales frequencies continuously; this remaps bin
indices, quantized to round(k · remap). A harmonic series at f, 2f, 3f…
survives integer remaps in recognizable form — remap 2 folds a harmonic
spectrum roughly an octave down — but at remap 1.37 the partials land on a
grid nature never drew: some merge, some vanish, spacings go irrational-ish.
The result reads as bells, metal, ghosts of the input. That in-between space
is the instrument. Sweep remap slowly across 1.0 and you can hear the sound
leave reality and come back.
Two practical corollaries:
remapjust above or below 1 (0.9–1.1) is the subtle zone — a detuned, phasey shadow of the input, cheaper than it sounds.remapwell below 1 stretches the low spectrum upward across the output (each output bin reads a lower input bin), thinning the top; well above 1 compresses everything into the bass and discards the input's top octaves entirely. Loud, dark, and blunt — usually wants a fresh brightness source afterwards.
The knobs
There is really one, plus the frame:
remap — the rule
Continuous. Identity at 1; integer values are the quasi-musical landmarks; everything between is the inharmonic wilderness. Automate it slowly — the per-frame quantization means fast sweeps step audibly, which is either the problem or the point.
FFT size — the grain of the grid
Bigger frames put the bins closer together, so the remap grid is finer: less
quantization grit, smoother inharmonicity, more latency (one frame, as
always) and more transient smearing. Smaller frames make the remap chunkier
and more overtly digital. Unlike tap.nr~, where the frame is a fidelity
question, here it is a flavor question.
Recipes
- Bell foundry: harmonic material (a
tap.vco~saw, a piano) atremap 1.3–1.6, into a long reverb. Instant inharmonic percussion. - The shadow voice: speech at
remap 0.95, mixed subtly under the dry — a wrongness the ear notices before the mind does. - The corridor: automate
remap1.0 → 2.0 over a minute under a sustained chord — a slow departure from consonance that lands, at exactly 2, somewhere almost stable again. - Stacked plumbing: two in series at
remapa and b is a remap at a·b with two layers of quantization grit — the grit is the reason to do it.
When it is not the right tool
- Musical transposition. The remap is spectral plumbing, not pitch
shifting: use
tap.shift~for clean intervals,tap.pitchaccum~for the spiral. - Harmonizing or formant work. Nothing here knows what a formant is; the rule moves bins, not vowels.
- Subtle timbre correction. Even at its gentlest this object is a
character effect; EQ-shaped intentions belong with
tap.filter~ortap.svf~'s EQ modes.
Checkpoint
One rule — output bin k reads input bin round(k · remap), Hermitian-mirrored — inside a transparent STFT: identity at 1 (pinned), octave-fold at 2 (pinned), and an inharmonic wilderness everywhere between. The FFT size sets the grain of the grid, the sweet spots are narrow, and that is the appeal: this is the book's one unapologetic reality-distortion tool. Use it where nature's spectra have gotten boring.
The acid machine
The Roland TB-303 was designed to imitate a bass guitar, failed completely,
and accidentally defined thirty years of dance music. What makes it
unmistakable is not any one block — a saw into a lowpass is every synth ever
made — but the coupling: accent drives the filter and the amplifier through
shared circuitry with memory across notes, slide is a gate that refuses to
let go, and the envelopes are fixed RC discharge curves with exactly one knob
between them. tap.303~ is a circuit-informed model of that whole tangle;
tap.diode~ is its filter as a standalone object; tap.303.seq~ is the
other half of the instrument. This chapter is what each control trades, and
what the measurements say the model actually delivers.
Companion material: the reference pages and help patchers in the TapTools-Max
package, and two executed verification notebooks —
tb303.ipynb
for the voice and
step_seq.ipynb
for the sequencer — every number below is a measurement from one of them or
from the kernel test suite. Provenance runs through Tim Stinchcombe's filter
analysis, Robin Whittle's Devil Fish documentation, the x0xb0x schematics,
and Robin Schmidt's Open303, whose measured calibrations several constants
adopt verbatim.
What the hardware is, in one paragraph
One saw-core oscillator (the "square" is the saw through a transistor
shaper, not a clean pulse), into a four-stage diode-ladder filter — not
the Moog transistor ladder; the diode ladder's stages load each other, which
is why its resonance is broader, less pure, and entirely its own — then a
one-transistor amplifier. Two envelopes, both decay-only RC discharges: the
Main Envelope sweeps the cutoff (the envmod knob decides how much), the
VCA envelope is fixed. Accent makes the Main Envelope hotter and faster,
routes it into the VCA, and charges a capacitor (C13) through the resonance
pot — and because C13 doesn't fully discharge between closely spaced accents,
runs of accented notes bloom, the famous wow. Slide holds the gate across
the step boundary while the pitch CV glides through a ~60 ms RC. Everything
about a note — pitch, gate, accent, slide — comes from the sequencer, not
the panel. That is why this is three objects, not one.
The filter first: tap.diode~
The panel says "18 dB/oct"; the circuit is four poles whose asymptotic slope is 24 dB/oct with a shallower region near cutoff — Stinchcombe untangled this, and the kernel reproduces his published transfer function to 0.028 dB. Two behaviors are load-bearing and easy to get wrong:
- The resonance feedback runs through a 150 Hz high-pass, so resonance thins as the cutoff drops — low notes squelch, they don't ring. Pinned by test: the ring-down Q falls with cutoff.
- A stock 303 never quite self-oscillates, and neither does this filter
at stock settings. That emerged from the modeled feedback high-pass rather
than being programmed in, and it's documented as a trait, not a defect.
(Push
resonancepast 1.0 — the bend range runs to 1.5 — and it will sing for you anyway.)
Like tap.ladder~ it has a solver choice: fast (default) or exact,
which iterates the re-linearized solve to convergence on the true nonlinear
loop. Measured across a matrix out to
resonance 1.4 and +24 dB drive — beyond anything the hardware can reach —
the two differ by at most −44.9 dBr, at 1.6–3.3× the CPU. The exact
solver is there for the suspicious; the fast one is there for the patch.
oversample (1/2/4, default 2) and a signal-rate cutoff in the right inlet
round out the tap.ladder~ surface.
The voice: tap.303~, knob by knob
The attributes mirror the seven-knob panel; the calibrations are Open303's measured laws.
The blocks are ordinary; the red and amber wires are the 303. Accent touches three destinations at once, and C13 remembers across notes — the couplings are the instrument.
waveform—saworsquare. The square is the hardware's shaped saw:−tanh(10^(36.9/20)·saw + 4.37), Open303's measured constants verbatim — rounded and notched, audibly not a 50 % pulse.cutoff— the knob in Hz. Stock travel is the measured 302–2394 Hz; the attribute range (100–5000) is a flagged bend beyond the panel.resonance— 0..1 is stock; up to 1.5 is the bend.envmod— how much Main Envelope reaches the cutoff, with the hardware's measured law: 2/3 of the sweep goes above the knob position, 1/3 below, and the "gimmick" offset shifts the resting point down as you turn it up. The knobs feel right because the interaction is modeled, not just the ranges.decay— Main Envelope decay, 200 ms–2 s. On an accented note the hardware ignores this knob and runs at ~200 ms; so does the model (adjustable via theaccdecaybend, 50–2000 ms).accent— how hard accented notes hit: louder and punchier (the envelope routing), and quackier (the C13 sweep, scaled by the resonance knob). The wow is measured: over a run of closely spaced accents the cutoff peak builds by ×1.94, and decays back within ×0.998 once the accents stop. Consecutive accents at high resonance are the entire genre.tuning,gain— cents and dB. Plumbing.
The envelopes carry the schematic's fixed interrelations: MEG attack ~3 ms,
VCA attack ~3 ms with a measured ~1.23 s decay chopped at gate-off, 50 ms when accented.
None of these have knobs on the hardware, so none of them have knobs here —
except through the documented Devil-Fish-style bends (slide 10–500 ms,
attack 0.3–30 ms, accdecay, and drive ±24 dB into the ladder, where
the diodes compress: +24 dB of gain buys only 9.2× of RMS). All stock at
their defaults.
Phase 2 added vca clean|warm: the one-transistor class-A stage as a
slope-normalized biased saturator, in the hardware's signal order. The
distortion tracks the envelope — measured 5.4 % difference signal on quiet
notes, 11.5 % on hot accents — so warm thickens exactly where the hardware
does. clean (default) is bit-identical to phase 1.
House machinery throughout: seed/tolerance per-unit component spread (an
mc. stack of 303s with different seeds detunes and drifts like a wall of
real units), 16 preset-morph slots with factory acid in 1–8
(squelch, sub, screamer, rubber, knock, bloom, overdriven, glass), and
per-sample ramps on every parameter.
The note interface, and why slide is free
tap.303~ is TapTools' first pitched instrument, and its inlets are the
package-wide melodic contract: pitch as a MIDI note number signal in the
left inlet, gate with amplitude-as-accent in the right — 1.0 is a plain
note, 2.0 fully accented (depth = amplitude − 1). Slide needs no input at
all: a pitch change while the gate is held is a slide — legato, no
envelope retrigger, the ~60 ms RC glide — which is exactly the hardware's
own definition. A note <pitch> [accent] [slide] message covers patching
without signals.
The other half: tap.303.seq~
Half the 303's sound is sequencer behavior, so the sequencer emits the
voice's contract verbatim: a pitch signal and a gate signal, clocked by a
phase ramp (0..1 per pattern, a phasor~). Per step: pitch, gate/rest,
accent, slide. The measured facts, from the sequencer notebook:
- Steps land on the analytic grid within one sample; the gate opens at
the step start and closes at 0.5 of the step (Open303's
stepLength). - A slid step is approached with the gate held: 16 gated steps with 3 slide flags produce exactly 13 note-ons — the other three arrive legato, pitch stepping on the boundary sample, and the voice glides.
- Accented steps gate at 2.0;
transposeshifts live, like the hardware's transpose mode without the mode;swingand pattern slots with cycle-quantizedrecallare shared with the drum rows (next chapter).
The 1981 pitch-mode/time-mode data entry is deliberately not recreated. You keep the data model; you lose the part everyone hated.
When it is not the right tool
- You want a generic bass synth.
tap.vco~+tap.svf~+tap.adsr~give you ADSRs, waveform variety, and a filter that behaves. This object's value is its refusal to decouple. - You want the filter without the biography —
tap.diode~alone, ortap.ladder~if you want the Moog character instead of the 303's. - You want polyphony. It's a monosynth;
mc.gives you many monosynths, which is not the same thing as a polysynth and shouldn't be.
Checkpoint
A diode ladder that matches the published analysis to 0.028 dB and won't self-oscillate until you bend it; a voice whose envelopes, accent path, and C13 memory come from the schematic, with the wow measured at ×1.94 across an accent run; slide as pure gate-hold, so legato falls out of the note contract; and a sequencer that emits that contract sample-accurately. The coupling is the instrument — and every claim above has an executed notebook cell behind it.
The drum machine
The Roland TR-808 is the most thoroughly analyzed drum machine in the
academic literature, and the reason is charming: the whole instrument is
analog synthesis. No samples anywhere — every sound is a small circuit, and
most of them are variations on about four ideas. The tap.808.* family
recreates the eight voice channels circuit block by circuit block, one
external per channel, and tap.808.seq~ supplies the machine's other half
as one sequencer row per patch cord. This chapter is the family tour: the
shared trigger contract, each voice's character and knobs, and the
calibration pass against a real unit that the numbers come from.
Companion material: each voice's reference page and help patcher, the family
overview patcher (tap.808.maxhelp — all eight voices sequenced off one
phasor~), the
tr808_calibration.ipynb
notebook, and the
step_seq.ipynb
sequencer notebook. Provenance runs through the Werner–Abel–Smith papers
(DAFx-14 and companions) and the TR-808 Service Notes, read component by
component; every magic constant in the kernel headers carries its schematic
designator.
One trigger to rule them all
On the hardware, every voice hangs off a common trigger bus: the CPU's 1 ms
pulse rides a voltage between 4 and 14 V depending on the accent circuit,
and a hotter pulse excites each circuit harder — more punch, slightly
different timbre — not merely louder. The family keeps that literally:
every voice fires on a signal rising edge, and the edge's amplitude
(0..1) is the accent, mapped onto the 4–14 V bus. bang and
trigger 0.7 messages cover the scheduler side. Filter states persist
across triggers, so fast rolls interfere with the ringing tail like the
hardware — no machine-gun effect. And because the excitation is a voltage,
anything that makes an edge can play the kit: a click~, an envelope, a
tap.303.seq~ gate, or the row object built for the job.
The voices
tap.808.kick~ — the bridged-T with a biography
The bass drum is a damped bridged-T resonator (~49.4 Hz from the modeled
component values; Roland's chart optimistically says 56, real units measure
as low as 48) with three behaviors that make it the kick, all emergent
from the modeled schematic: for the first ~6 ms the envelope saturates Q43
and the resonator sits near ~129 Hz — the attack punch, which is a
different mechanism from the famous downward pitch "sigh" (leakage through
R161, the paper's fitted nonlinearity); and a retriggering pulse re-excites
the center node as the envelope collapses so the note doesn't step down.
Panel knobs: decay (seconds of ring at the top), tone (click at ~7 kHz
down to ~300 Hz), level. Paper-documented bends, stock by default:
tuning, pulse, sigh, attack — turn sigh 0 and the pitch relaxation
disconnects, exactly as the bend does on the bench.
Calibration: against a real unit's knob-gridded sample set, the fundamental sat within 2.4 % at every tone/decay position and the −40 dB decay endpoints within 6 % (72 ms → 2.36 s measured, 69 ms → 2.42 s modeled) — no constant needed changing.
tap.808.snare~ and tap.808.clap~ — resonators plus noise
The snare is two bridged-Ts (the late-revision ~173/336 Hz pair) with a
trigger divider and the "snappy" path — enveloped noise, band-limited around
4 kHz to the measured unit. Fundamentals calibrated within 1.2 %, including
the mode flip at tone-max. The clap channel (@model clap|maracas) is the
Service Notes' Figure-13 circuit: band-passed noise near 2 kHz through a VCA
driven by a three-teeth sawtooth retrigger — the "multiple hands" transient —
plus the Q70 reverberation tail. The maracas mode is the same noise voiced
short and bright.
tap.808.hat~ and tap.808.cymbal~ — the metal bank
Six Schmitt-trigger square oscillators (205.3, 369.6, 304.4, 522.7 Hz plus
the two trimmer-tuned at 800 and 540, duty 47.98 %) feed two bandpass
voicings near 3.4 and 7.1 kHz. Werner et al. measured that resistor variance
puts any given unit up to ~20 % off those frequencies — which is why no
two 808s' cymbals sound alike, and why seed/tolerance exists: every seed
is a different unit off the line, and an mc. stack of cymbals decorrelates
like real hardware. The hats are one object with two trigger inlets
because on hardware they are one circuit with two envelope paths and a
choke — closed chokes open (the Q23/R173 path), pinned by test, and
unimplementable as separate externals. Open-hat decay spans the chart's
90–600 ms; the cymbal's two separately enveloped bands cover its 350–1200 ms
"sizzle" span. tap.808.cowbell~ taps just the 540/800 pair into the ~860 Hz
voicing with a two-slope envelope; more cowbell is a patching decision.
tap.808.tom~ and tap.808.rim~ — the resonator variations
Six sounds on two objects, as the hardware switches them: @size low|mid|high
× @model tom|conga. Congas are the tom circuit without its noise layer,
tuned differently; the toms add the D80/D81 attack pitch fall and a pink
noise layer. The rim channel is @model rimshot|claves: the rimshot's
~1667 + 455 Hz crack with the swing-VCA's harmonics, versus the claves' pure
~2500 Hz tick. Tunings sit within ~4 % of the measured unit.
Roland's universal voice circuit. Eight voices, one network — the kick earns its punch by modulating the leg per sample.
The calibration pass, honestly
The §7.2 calibration ran against a real TR-808 (s/n 103852) recorded from the individual outs with knob positions encoded in the filenames — a 0/2.5/5/7.5/10 dial grid, 116 samples — which upgraded "sounds right" to a quantitative per-knob-cell comparison. Identical measurements (spectral-peak fundamental, −40 dB decay, power centroid) ran on both sides. The pitches were already right nearly everywhere; what the pass actually changed was time: tom, conga, cowbell, and clap tails roughly doubled to match the unit, the snappy was band-limited and re-enveloped, the rimshot re-voiced low-dominant, the cymbal's decay span corrected. Each kernel header carries its residuals. The lesson generalizes: schematics get you the frequencies; recordings get you the envelopes.
The other half: tap.808.seq~
One row of the 16-step sequencer, as an object: feed it a phase ramp (0..1
per pattern, a phasor~) and it emits trigger impulses whose amplitude is
the step's accent — the family contract, straight into any voice. Twelve
rows off one phasor are the hardware's panel, sample-locked forever; the
accent row falls out of giving every row the same accents list. The
measured facts, from the sequencer notebook: steps land on the analytic grid
within one sample; the pinned levels are plain 0.01 (the 4 V base — an
un-accented hit still strikes the circuit) and accented 0.5 (the accent
knob at noon; 1.0 is the full 14 V); swing delays the off-16ths by exactly
swing/2 of a step; a length 12 row against 16s is the triplet pre-scale
generalized to polymeter; and pattern slots with cycle-quantized recall
are the A/B-half and fill switching as one message. pulse widens the
impulse into a held gate when you'd rather drive tap.adsr~ than a drum.
When it is not the right tool
- You want a kick, not the kick. A sine with an envelope is cheaper and takes EQ more politely. This family's value is the circuit behavior — the attack jump, the choke, the accent-as-voltage.
- You want your own drum sounds. These circuits are what they are;
seed, the documented bends, and the panel knobs bend them, but a sampler is a sampler. - You want 909 hats. The 909's metal is sampled; this machine's is six square waves. Different instrument, different chapter, maybe someday.
Checkpoint
Eight channels, four circuit ideas — bridged-T resonators, a shared metal bank, noise paths, swing-VCAs — under one amplitude-as-accent trigger bus, calibrated per knob cell against a real unit and honest about what changed (the tails) and what didn't (the tunings). The hats choke because they share a circuit; the cymbals decorrelate because resistors do; and the sequencer row emits the same voltage idea the voices drink, so the whole kit runs off one phasor ramp. The machine's two halves, both measured.
The note you meant
Every sung note is two notes: the one that happened and the one you meant.
tap.tune~ measures the distance between them and closes it — how fast it
closes it is the whole instrument. Closed slowly, nobody knows it was there.
Closed instantly, everybody knows: that snap is the most famous vocal
effect of the last twenty-five years. One object, one time constant, both
worlds.
A short history matters here, told plainly. The classic pipeline — detect
the pitch, snap it to the nearest allowed note, retune by time-domain
resynthesis — was patented in 1998 and the patent expired in 2018, which is
why a whole field of tuners exists today and why this object can implement
the technique from the literature. The famous product name remains a live
trademark, which is why this object is called tap.tune~ and this chapter
says "hard snap" instead. And editing individual notes inside a chord
remains patent-fenced territory — tap.tune~ is monophonic by design, not
by omission. (None of this paragraph is legal advice; the project's own
ship-gate is a freedom-to-operate review.)
Companion material: the reference page and help patcher in the TapTools-Max
package, the runtime maxtest, and two executed notebooks — tune.ipynb
here and pitchshift.ipynb in the DspTap repo — that measured every claim
below.
A per-hop brain over a per-sample corrector, with one seam where three resynthesis engines interchange.
The knob that is the instrument: speed
speed is the time constant, in milliseconds, of the glide onto the target
note.
- 0 ms — the hard snap. The correction lands within a detection hop (~5 ms). Vibrato gets quantized into terraces; note transitions become instant staircase steps. This is the effect, worn on the outside.
- 10–40 ms — classic correction. Fast enough that a listener hears "a singer with good intonation," slow enough that the attack of each note — where identity lives — is not robotic. The default is 20.
- 100 ms and up — intonation leaning. The corrector arrives so late it only tames drift; vibrato passes through nearly untouched.
The notebook's pitch-track figure shows all three glides onto the same
46-cent-sharp note; the kernel test pins the exponential's arrival. There is
also amount (0–100%): a fader on the correction distance itself. 100
lands on the target; 50 splits the difference — a gentler kind of honesty
that keeps a performance's shape while shrinking its errors.
Telling it what is allowed
The corrector never invents a target; it snaps to the nearest note you allowed.
key+scale— the usual contract:@key d @scale majorand every detected pitch pulls toward the nearest D-major degree. Presets: chromatic, major, minor, harmonic, melodic, pentatonic, minorpentatonic.notes— the twelve toggles, absolute pitch classes C through B, panel-style:notes 1 0 0 0 1 0 0 1 0 0 0 0snaps everything to a C-major triad, which is less a correction than an arrangement decision.mode midi— the target is the nearest currently held MIDI note (note 64 100holds E4; velocity 0 releases;flushclears). Hold one note and everything becomes that note; hold a changing chord's roots and the corrector is suddenly a performable melody-mangler. No notes held means no correction — the object never guesses.
An empty mask behaves the same way: nothing allowed, nothing changed.
Three engines, one corrector: backend
Detection, targeting, and the glide are shared; only the resynthesis swaps. All three land the same intonation — the notebook drives the same vibrato "voice" through each and all three settle on 220.00 Hz — so the choice is about character and latency, not accuracy.
| backend | what it is | choose it for | latency @ 48 kHz |
|---|---|---|---|
grain | two-tap delay-line, window locked to the detected period (the tap.shift~ engine) | the default; lowest latency, waveform-preserving, happy on any material | a few ms |
psola | true TD-PSOLA | voice — it preserves formants by construction | ~36 ms |
pvoc | peak-locked phase vocoder | dense, harmonically rich material; pairs with formant | ~21 ms |
Switching live is click-safe: the incoming engine starts from silence and
fades in rather than splicing stale audio. One honest caveat per engine:
grain colors sustained unpitched input with a mild moving comb (the
known trade of its class); psola wants harmonic material — on a pure
sine shifted far, its output legitimately thins (the machine chapter
explains why that is the same property as its formant preservation);
pvoc smears sharp transients slightly, as every phase vocoder does.
Keeping the singer's mouth: formant
A correction of thirty cents moves formants thirty cents — nobody hears it.
A MIDI-mode command of five semitones moves them five semitones — everybody
hears it; that is the chipmunk. @formant 1 enables LPC formant
preservation on the pvoc backend: the pitch moves, the vocal tract's
envelope stays where the singer put it. The notebook corrects a synthetic
voice up 5.5 semitones both ways; with the flag on, the formant bump stays
put (band-energy ratio 730:1 in its favor). psola needs no flag — formant
preservation is its resampling rule — and grain ignores the flag.
Letting it find the key: autokey
@autokey 1 starts a learner: every voiced detection drops its pitch class
into a histogram that forgets with about a minute of memory, scored against
the published Krumhansl–Kessler key profiles. Two design decisions worth
knowing:
- It never acts on its own. A key estimate that silently re-aimed your
targets mid-phrase would be a bug wearing a feature's clothes.
getkeyasks (the right outlet answerskey d major 0.95, orkey nonein the first half-second);applykeyadopts the estimate into thekeyandscaleattributes — visibly, where you can see and undo it. - It forgets on purpose. The one-minute memory means a modulation stops arguing with the old verse about as fast as you stop playing it.
The kernel test plays a D-major scale and reads back D major at 0.95 confidence; an A harmonic-minor melody reads as A minor.
The right outlet
While the input is voiced, the right outlet reports
pitch <midi> <hz> every @interval milliseconds (default 50; 0 disables;
a pitch -1 0 marks the end of voicing). That is a free tuner display, a
melody recorder, or the control signal for whatever you want to drive with
the singer's pitch — and it is the same detector the corrector itself uses,
so what you see is what it acted on.
Recipes
- Invisible repair:
@scale major @key(your key)@speed 25 @amount 80. The 80 keeps a little humanity in the intonation; nobody will name what changed. - The famous one:
@speed 0 @scale minorpentatonic. Fewer allowed notes make the terraces wider and the snap prouder. Add melisma. - One-note choir:
@mode midi @speed 5, hold a note, feed it speech. Everything becomes chant on that pitch. - Formant-true transposer:
@mode midi @backend pvoc @formant 1 @speed 10, play a melody against a held vocal — a harmonizer that keeps the singer's identity. - Tuner display only:
@amount 0 @interval 20— the object corrects nothing and the right outlet becomes a clean pitch stream.
When it is not the right tool
- Chords. The detector is monophonic; a chord reads as garbage or as its loudest note, and per-note polyphonic editing is deliberately out of scope (see the history paragraph). Split voices first, or don't.
- Drums, breath, speech consonants. Unpitched input passes through with no correction — by design — but the grain engine adds its mild comb coloration to sustained noise. For processing unpitched material there are better rooms in this house.
- Creative shifting. If the goal is an interval rather than
intonation,
tap.shift~is the plain shifter,tap.harmony~the formant-preserving chord stack, andtap.pitchaccum~the spiral;tap.tune~always measures first and that measurement is latency you don't need.
Checkpoint
Detect, snap to the nearest allowed note, glide at speed — that is the
whole machine, and speed is the dial between honesty and effect. Targets
come from key + scale, twelve toggles, or held MIDI notes; three resynthesis
engines trade character against latency while landing the same intonation;
formant keeps the singer's mouth in place when corrections get big;
autokey learns the key but only ever suggests. The right outlet tells you
what it heard. And when the input isn't a single pitched voice, the honest
move — which the object makes — is to change nothing.
Distortion with a memory
Every distortion plugin can bend a transfer curve. tap.overdrive~ is built
on the observation that the pedals people actually love — the Tube Screamer
lineage, and specifically the Mad Professor Little Green Wonder that served as
this object's listening reference — don't apply one curve to the whole
spectrum. Their clipper lives inside an op-amp's feedback loop with
frequency-dependent parts around it, and that loop is most of the sound: bass
sees less gain and stays tight, mids break up first, and the knee never quite
flattens because the clean signal always rides through. A memoryless
waveshaper — including both modes of the Jamoma-era tap.overdrive~ this
object succeeds — structurally cannot do any of that. This one can, because
the shaper sits inside a lowpass feedback loop: distortion with a memory.
Companion material: the reference page and help patcher in the TapTools-Max
package, and the verification notebook,
where every number below is an executed, plotted measurement of the shipping
kernel. The figures in this chapter are measurements too — regenerated from the
same kernel through the C ABI by book/figures/overdrive.py, never drawn by
hand.
What the loop buys
The claim worth leading with, because no static curve can make it: the
object's small-signal gain tilts with frequency, and the tilt grows with
drive. Measured between 80 Hz and 4 kHz, the tilt is +5 dB at drive 0
(just the voicing EQ), +16.3 dB at drive 0.5, +17.2 dB at drive 0.9. Low
frequencies are pinned near-clean by the feedback while mids and highs take
the full drive gain — so a low E stays articulate under the same setting that
saturates the pick attack. That is the Tube Screamer "tightness" in one plot:
The measured headline. A memoryless shaper's version of this figure is three horizontal lines.
The second structural trait: the transfer never flattens. A unity clean path is summed around the clipper — the non-inverting op-amp topology — so however hard the shaped part saturates, output keeps rising with input (measured strictly monotonic at every drive setting). The old sine-shaper mode's hard ±1 plateau, a large part of what read as "digital," is gone by construction.
Compression without a ceiling: the slope falls as drive rises, but never to zero.
The knobs, one by one
drive — 0 to 1, edge-of-breakup to saturated
Normalized, like every musical parameter on this object, with the perceptual
mapping done inside (the knob sweeps the clipper's gain from +6 to +46 dB,
with a level compensation tracking it). drive 0 is a pedal's gain knob at
full counterclockwise — still warm, not bit-clean; bypass is the clean
switch. The normalized range maps directly onto MIDI/OSC controllers, and
onto Q15/Q31 fixed-point for the embedded ports this kernel is written to
survive.
body — the signature voicing control
The LGW's defining knob, reproduced as linear pre/post EQ around the clipper (that's what it is in the pedal — voicing, not nonlinearity). Toward −1, fuller lows reach the clipper and the top gets a slight shelf lift; toward +1, the lows thin and tighten and an upper-mid bell pushes forward — centered at 1150 Hz, deliberately above the classic TS hump. Measured at the extremes: 100 Hz moves by 10 dB, the 1150 Hz push adds 4 dB, the counterclockwise treble lift is +2.5 dB at 8 kHz. The exact centers and gains are by-ear placeholders pending the in-Max voicing pass against LGW demos — the shape of the control is final, the seasoning isn't.
The knob's whole range. Note the crossover around 500 Hz: body trades lows
against upper mids around a stable center, like the pedal.
asymmetry — the even harmonics the old object couldn't make
Both Jamoma modes were odd functions: odd harmonics only, the entire "warmth"
vocabulary absent. asymmetry biases the clipper: at 0 the path is exactly
symmetric (measured H2 at −151 dB — the numerical floor), and raising it
brings the even series up smoothly (H2 at −26 dB by asymmetry 0.6). The
default sits at 0.15, a small nonzero warmth chosen by ear. Asymmetric
clipping generates DC, so a DC blocker sits permanently after the clipper —
measured output mean under full drive, full asymmetry: 10⁻¹⁰. (The original
TTOverdrive contained a DC blocker whose output was computed and then
discarded; this one is load-bearing.)
The same tone, the same drive — the only change is asymmetry, and the even
series (H2, H4, …) appears between the odd lines.
oversample — 1, 2, 4, or 8; default 4
Clipping makes harmonics; harmonics past Nyquist fold back as inharmonic junk. At 1× a hard-driven 5 kHz tone puts its folded seventh harmonic at −22 dB relative to the fundamental — clearly audible garbage at 12993 Hz. At the default 4× the same component measures −36 dB, with the true harmonics unchanged. Turn it down to 1× only when CPU matters more than the top octave, or when you want the fizz.
preamp, output, smooth, bypass, mute
Input and makeup gain in dB (±24) — the only unit-bearing parameters, because
gains are the one place real units belong. Everything ramps click-free over
smooth milliseconds (default 20).
Where it sits in a patch
Mono by design; wrap it in mc. for multichannel like the rest of the
package. It takes line-level signals as happily as guitar DI — the drive
mapping is normalized to full-scale digital, not to pickup output. For the
LGW move, start at drive 0.4, body -0.3, asymmetry 0.15 and ride body
against the source's low end. For a clean boost that just thickens, drive 0
with asymmetry 0.3. For fuzz territory this is the wrong object on
purpose — the loop keeps pulling it back toward articulation.
Every claim above is pinned twice: as an executed measurement in the
notebook, and as a hard assertion in the kernel's Catch2 suite
(tests/overdrive_test.cpp), which CI runs on every push. The math behind
the loop — including why it had to be solved zero-delay, and what happens if
you don't — is in the machine chapter:
The clipper in the loop.
Solving the filter on paper: svf.h
The user-facing chapter promised that tap.svf~ is "unconditionally stable
under per-sample cutoff modulation" and that its morph corners are
"bit-identical to the discrete modes." Promises like that are either
mathematical facts or marketing. This appendix does the math: it derives the
filter the way the file was actually designed, then walks the engineering
decisions that don't show up in a Bode plot — and why each one beat its
alternative.
The reference is Andy Simper's Cytomic technical papers ("Solving the
continuous SVF equations using trapezoidal integration and equivalent
currents"), specifically the SvfLinearTrapOptimised2 form. What follows is
the same derivation with the file's variable names.
The analog prototype, and where digital versions go wrong
The state-variable filter is two integrators in a loop. With cutoff ω and damping k = 1/Q, the continuous equations are:
v1' = ω · (v0 − k·v1 − v2) (band state: input minus damping minus low state)
v2' = ω · v1 (low state: integral of band)
Lowpass is v2, bandpass v1, highpass v0 − k·v1 − v2 — every response lives in the same two states, which is what makes an output mix (and therefore a morph) possible at all.
The textbook digital version (Chamberlin) discretizes with explicit Euler: each integrator uses the previous sample's value. That inserts a unit delay into the loop, and a delay in a feedback loop is a stability bomb with a frequency fuse: the design blows up as fc approaches fs/6, and modulating the cutoff re-lights the fuse every sample. The classic workarounds (oversample it, clamp it) treat symptoms.
Trapezoidal integration, and the algebraic loop
The fix is to integrate with the trapezoidal rule — average the old and new derivative — which in filter terms is the bilinear transform. Define the prewarped gain the file computes once per cutoff change:
g = tan(π · fc / fs)
(The tan is the prewarp: it makes the digital filter's response at fc exactly match the analog prototype's, all the way to Nyquist. This single line is why self-oscillation later measures 999.7 Hz for a 1 kHz setting rather than drifting flat.)
Trapezoidal integration of state s with input x is s_new = s + g·(x_old + x_new). Grouping the "old" terms into a memory variable — Simper's equivalent current ic = s + g·x_old — each integrator becomes:
s_new = ic + g · x_new with ic updated as ic_new = 2·s_new − ic
But notice the trap: x_new for the first integrator is the new band value, which depends on the new low value, which depends on the new band value. The new sample appears on both sides — an algebraic loop, exactly the "zero-delay feedback" the initials ZDF refer to. Instead of breaking the loop with a delay (Chamberlin's sin), we solve it. It is linear, so substitution gives a closed form. With v0 the input and ic1, ic2 the two equivalent currents:
v1 = a1 · ic1 + a2 · (v0 − ic2) the band state, solved
v2 = ic2 + g · v1 the low state, then follows
where a1 = 1 / (1 + g·(g + k)), a2 = g · a1
Those are precisely the file's per-section solve constants (a1, a2, a3 = g·a2 caches the product used by the low state), and the state update is the
canonical TPT pair ic1 = 2·v1 − ic1, ic2 = 2·v2 − ic2.
Why this is unconditionally stable, even modulated: the trapezoidal rule is A-stable — it maps the entire left half of the s-plane (every stable analog filter) inside the unit circle, for any g > 0. Change g every sample and each sample still computes a passive, energy-consistent step; there is no regime of fc or modulation rate where the update gains exceed unity. The notebook's 90 Hz-LFO-through-five-octaves torture test isn't surviving by margin; it's surviving by theorem.
The loop the algebra just solved, and the mixer the next section explains.
The output mix, and why morph corners cost nothing
Every response is a weighted sum over the same solved values:
y = m0·v0 + m1·v1 + m2·v2
lowpass m = (0, 0, 1) notch m = (1, −k, 0)
bandpass m = (0, 1, 0) peak m = (1, −k, −2)
highpass m = (1, −k, −1) allpass m = (1, −2k, 0)
mode_morph linearly interpolates the mix vector around the circle LP → BP →
HP → notch → LP. Two facts follow by construction, not by tuning:
- At a corner, the interpolated vector equals the discrete mode's vector exactly — same floats, same states, same arithmetic. The notebook's measured max difference of 0 is not a tight tolerance; it is an identity.
- Morphing is free. The states don't know the mix exists; sweeping it can never destabilize anything, because it is three multiplies downstream of the filter.
The parametric-EQ trio (bell, shelves) is the same machinery with mix weights that depend on a gain factor A = 10^(dB/40), straight from Simper's tables — and it always runs a single section, because cascading an EQ stage squares its boost: two +12 dB bells are a +24 dB bell, which is never what the user typed.
The cascade: Butterworth spread, resonance on the last section
Orders 4 and 8 run two and four sections at the same cutoff. Stacking identical Q = 0.707 sections would droop the passband (each contributes its −3 dB early); instead the sections take the Butterworth Q spread — the Qs whose product of section responses is maximally flat:
Q_i = 1 / (2·cos θ_i), θ_i the Butterworth pole angles
order 4: 0.5412, 1.3066 order 8: 0.5098, 0.6013, 0.9000, 2.5629
That is why the measured response sits at −3.01 dB at fc at every order. User resonance then sharpens only the final (highest-Q) section, via
Q_res = Q_base / (1 − r), r ∈ [0, 1) (clamped at 1 − 10⁻⁴)
— one clean resonant peak riding a flat passband, rather than four peaks
compounding. The inverse mapping (resonance_from_q) exists so the wrapper's
q message round-trips exactly.
The driven circuit: one saturation, one pass
The driven circuit places tanh on the band node — in the damping path, where an OTA's transconductance actually compresses. That placement is the whole design: as amplitude grows, the effective damping k·tanh(v1)/v1 grows with it, which is an automatic gain control wrapped around the resonance. Push the loop gain slightly past the oscillation threshold at resonance 1.0 and the filter must oscillate (the linear model's poles are outside the circle) but cannot run away (the saturation restores effective damping as amplitude rises). Bounded self-oscillation is not a limiter bolted on; it is the fixed point of that tug-of-war. An all-zero state solves the equations too — hence "give it a ping."
Solving a nonlinear zero-delay loop exactly needs iteration. The file uses
the one-pass scheme shared with tap.ladder~'s solver_fast: solve the
linear ZDF prediction for the band node, saturate it, commit. The error of
that shortcut is second-order in how much tanh bends over one oversampled
step — and the driven circuit always runs oversampled (2× default), with
4th-order Butterworth anti-image/anti-alias biquad pairs on the way up and
down. At these rates the one-pass and iterated answers are audibly identical;
the ladder file, which drives its nonlinearity much harder, is the one that
also ships a Newton option.
The engineering ledger
Decisions visible only in the code, with their reasons:
- Two-tier coefficient update. Recomputing everything per sample costs a
tan()plus the mix logic even when nothing changed. The file splits state: a shape tier (damping, mix weights, EQ gains — dirtied only when a non-frequency parameter or mode changes) and a cutoff tier (tan and the three solve constants — recomputed only when the incoming cutoff differs from the cached one). Signal-rate modulation pays for exactly what it moves. The benchmark ratchet recorded the win: modulated 2nd-order lowpass 36 → 19 ns/sample, modulated morph 77 → 28, bit-identical output (the morph-corner identity tests pin that "bit-identical" is literal). ramp_todoesn't dirty the shape tier for frequency — the cutoff cache catches it. One branch, measurable at audio rates.- Multichannel by frame protocol. Coefficients are computed once per
tick()and shared by every channel'sprocess(ch, x)— an N-channel engine outside Max for the cost of one solve. The Max wrapper stays mono by house rule (mc.wraps it). - Allocation discipline. The only allocation is the per-channel state
vector in
prepare(); setters are wait-free and safe from the message thread while audio runs, because a "set" is a ramp target plus a dirty flag. - Anti-denormal guard on the states (the
tap.comb~idiom): a filter ringing out into silence otherwise wanders into denormal territory and multiplies its own CPU cost right when the music is quietest. - What is deliberately absent: fast-tanh approximations and a polyphase halfband resampler are both flagged in the file as candidates — and parked, because each changes output microscopically and the project's rule is that optimizations land only bit-identical or explicitly signed off.
Checkpoint
Trapezoidal integration turns the SVF's two integrators into a solvable linear system per sample — A-stability is where the modulation-proofness comes from, prewarping is where the tuning accuracy comes from, and the output mix is where morphing comes from, corner-exact by construction. Butterworth spread keeps cascades flat; resonance sharpens one section; the driven circuit's tanh placement makes bounded self-oscillation a fixed point rather than a feature. The rest is bookkeeping — and the bookkeeping was benchmarked.
The nonlinear loop: ladder.h
The user-facing chapter promised self-oscillation in tune (8009 Hz
measured for an 8 kHz cutoff), THD that walks from 0.5 % to 33 %, comp
recovering exactly the passband resonance eats, and a measured 13.5 dB
from oversampling. This appendix derives all of it. The
SVF appendix built the trapezoidal machinery; here it is wrapped
in four tanh saturators and a feedback loop supposed to go unstable —
and the engineering is solving a loop that no longer solves on paper.
The stage: one pole, trapezoidal, prewarped
Each stage is the analog one-pole lowpass y' = ω·(x − y), discretized with
the trapezoidal rule exactly as in the SVF. Once per cutoff change,
update_derived computes
g = tan(π · fc / fs_os) the prewarped integrator gain
m_g = g / (1 + g) the solved per-stage gain, G below
and the per-stage step (tpt) is the standard zero-delay one-pole:
v = (x − s) · G y = v + s s ← y + v (= 2y − s)
For a lone lowpass the tan prewarp is a nicety. Here it is the tuning
system: the filter's oscillation frequency is set by where the stages put
their phase, so pole mis-placement becomes pitch error — the failure of
the classic Stilson/Smith ladder in the top octaves.
The loop: why the magic number is four
Four stages in series, global negative feedback: the stage-1 input is
L − m_k·y4 (L the driven input, m_k = 4.0 * resonance). The 4 is the
linear loop's oscillation threshold. At the cutoff a one-pole has response
1/(1 + j): magnitude 1/√2, phase −45°. Four in series:
|H⁴(fc)| = (1/√2)⁴ = 1/4 ∠H⁴(fc) = 4 · (−45°) = −180°
The input subtraction supplies the other 180°, so at fc — and only there — the loop phase is 360°. Barkhausen: oscillation begins at unity loop gain,
k · 1/4 = 1 ⇒ k = 4
So resonance = 1.0 (k = 4) is the mathematical edge, oscillation happens
at the tuned cutoff, and self-oscillation frequency is the tuning test:
the notebook measures 1000.2 Hz for a 1 kHz cutoff (0.02 % error) and
8009.0 Hz for 8 kHz (0.11 %) — the prewarp holding at the top of the
keyboard, as promised.
The file allows k_res_max = 1.1, i.e. k = 4.4, "comfortably past
self-oscillation." Past the edge the linear model diverges — but as
amplitude grows, tanh's small-signal gain falls, the effective loop gain
sags back toward 4, and the oscillation parks where they balance: the SVF
driven circuit's fixed-point argument, no clipper needed. The notebook's
oscillation (resonance 1.08) peaks near |y| = 0.10; the kernel test holds
five seconds at k = 4.4 finite, under 2.0 peak, RMS steady within a
0.7–1.4× band. An all-zero state also solves the equations — hence the
header's advice to ping it.
The file as a schematic: the Barkhausen condition lives at the red tap.
The algebraic loop, solved linearly first
Zero-delay feedback through four stages means y4 depends on the stage-1 input, which depends on y4. Linearly, substitution closes it: chain the linear stage form y = G·x + B·s (B = 1 − G, s the held state) through all four with u = L − k·y4,
y1 = G·u + B·s1
y2 = G²·u + G·B·s1 + B·s2
y3 = G³·u + G²·B·s1 + G·B·s2 + B·s3
y4 = G⁴·u + G³·B·s1 + G²·B·s2 + G·B·s3 + B·s4
then name the state-only part S and solve:
S = G³·B·s1 + G²·B·s2 + G·B·s3 + B·s4
y4 = G⁴·(L − k·y4) + S ⇒ y4 = (G⁴·L + S) / (1 + k·G⁴)
That last expression is predict_linear verbatim — the code's
(G2*G2*L + S) / (1.0 + m_k*G2*G2) with G2 = G*G and the same four-term
S over m_s1..m_s4. For the linear ladder it is exact — the four-stage
analog of the SVF's a1/a2 solve.
The saturators, and the one-pass commit
With tanh in every stage, the honest loop equation y4 = F(L − k·y4) has no closed form. The file ships two answers.
solver_fast (default) is Huovilainen-flavored prediction-correction:
compute predict_linear(L) as if the saturators weren't there, then run
the saturating stages once with that feedback value and commit (core):
t0 = sat(L − m_k·y4_est)
y1 = tpt(m_s1, t0, G), y2 = tpt(m_s2, sat(y1), G), ... y4 likewise
The committed y4 is not the y4_est the feedback used — that mismatch is the method's error. When the signal is small, tanh is the identity and the prediction is exact: the linear filter is recovered in the limit. The error grows only with how far tanh bends over one sample's state change — so drive × resonance is the failure axis, and oversampling (2× default) doubles as accuracy: it shrinks the per-step change being predicted.
solver_exact solves the true loop: Newton iteration on
F(g) = y4_trial(L, g) − g, where y4_trial evaluates the four
saturating stages for a guessed feedback value without touching state.
Seeded by the linear prediction, clamped to ±3 (a tanh-bounded loop cannot
park a fixed point far outside ±1), with a numerical derivative that falls
back to the seed if it degenerates, at most 12 iterations to a 1e-12
residual. The commit reuses the same core path — with the converged g
it reproduces the trial values while tpt advances the states. One code
path, two accuracies.
How different are they? The kernel test renders both at drive 3 dB,
resonance 0.5 and pins the maximum sample difference below 0.01; the
stress test (drive 24 dB, resonance 1.1, asym 1.0) asks only that
solver_exact stay finite and bounded. Audibly identical until drive
and resonance are pushed — and solver_fast costs one saturated pass
where Newton can cost dozens of trial evaluations per sub-sample.
asym: moving the operating point
Real ladder transistors don't match, so real stages don't saturate symmetrically. The model is an operating-point shift in every stage:
m_sat_bias = 0.3 · asym
sat(v) = tanh(v + m_sat_bias) − m_sat_dc m_sat_dc = tanh(m_sat_bias)
The subtraction keeps sat(0) = 0 exactly — silence in, silence out. Expand
tanh about the bias: a curvature term −tanh(b)·sech²(b)·v² appears only
when b ≠ 0, and a v² term generates second harmonic and DC. The notebook
measures the driven 2nd harmonic at −155.8 dB relative to the fundamental
at asym 0 (numerical noise — tanh is odd), rising to −18.6 dB at asym 0.6.
The DC is the rectifying side of the same v² term; the header owns it
honestly and delegates to tap.dcblock~. Drive's own numbers: THD
0.54 / 3.45 / 16.50 / 33.07 % at 0 / 8 / 16 / 24 dB (notebook) — odd
harmonics only, until asym says so.
comp: the passband bargain, quantified
At DC every stage passes unity; the closed linear loop gives
y4(DC) = L − k·y4(DC) ⇒ y4(DC) = L / (1 + k)
Resonance eats the passband by exactly 1/(1+k). At resonance 0.9, k = 3.6:
predicted 20·log10(1/4.6) = −13.3 dB; the notebook measures −13.2 dB. The
compensation is a pre-gain (update_derived):
m_in_gain = 10^(drive/20) · (1 + comp·m_k)
At comp = 1 the input is multiplied by (1 + k) and DC gain returns to exactly unity — measured +0.0 dB — with a linear blend below. One honest note: the compensation multiplies the input before the saturators, so high comp at high resonance also leans harder on the tanh stages — like turning up the level into hardware; authentic, not a linear post-trim.
Pole mixing: the Xpander table
core returns a fixed weighted sum over the taps [t0, y1, y2, y3, y4]
(k_c_mix). In the linear small-signal limit each tap is a power of the
one-pole response H applied to u, so the mixes are polynomial algebra:
lp12: y2 = H²·u
hp12: t0 − 2·y1 + y2 = (1 − H)²·u weights {1,−2,1,0,0}
hp24: (1 − H)⁴ → {1,−4,6,−4,1} alternating binomial
bp12: 2·(y1 − y2) = 2·H(1−H)·u → {0,2,−2,0,0}
bp24: 4·H²(1−H)² → {0,0,4,−8,4}
The binomial rows are literally (1−H)ⁿ expanded. The bandpass factors are unity-gain normalizers: at fc, |H| = |1−H| = 1/√2 with phases ∓45°, so H(1−H) has magnitude 1/2 and phase 0 — the 2 (and 4 for its square) restore 0 dB at center. Measured high-side slopes: 23.4 dB/oct for lp24 (want 24), 11.7 for lp12 (want 12). Two caveats, both inherited from the analog original: under saturation the taps carry distortion products and the algebra is approximate (the header says so), and the feedback is always the full four-pole loop — a "12 dB" mode is a two-pole slope riding four-pole resonance, exactly as in an Xpander.
Oversampling: paying for tanh honestly
tanh generates harmonics without limit; above Nyquist they fold back
inharmonically. run is the classic chain (the tap.verb~ pattern,
self-contained per house rule): zero-stuff by the factor, scaling the
retained sample by m_os to preserve passband gain; 4th-order Butterworth
anti-imaging at 0.45 of the original Nyquist (fc_norm = 0.45/m_os); the
nonlinear core at the high rate; a matching anti-alias Butterworth;
decimation by keeping the last filtered sub-sample. Each Butterworth is
two RBJ biquads at the textbook Q pair 0.54119610 / 1.30656296 =
1/(2·cos(π/8)), 1/(2·cos(3π/8)). Measured on a hard-driven 5 kHz tone:
non-harmonic (alias) energy −30.1 dB at 1×, −43.7 dB at 4× — the promised
13.5 dB. The clamp fc ≤ 0.49·fs_os keeps 20 kHz legal at every factor.
The engineering ledger
- One derived tier, not two. Unlike the SVF's split shape/cutoff
caches,
update_derivedrecomputes everything whenever anything moves (m_derived_dirtystays set whilem_ramps_active > 0). Nearly every derived value reads several parameters (m_in_gain: drive, comp, and resonance) — a finer split buys little. - The signal-rate cutoff path re-dirties deliberately.
process(x, cutoff_hz)recomputes for the override, then setsm_derived_dirty = true— "the cached G belongs to the override, not the parameter" — so the message-rate path never serves a stale one. - Ramps everywhere, counted. Every parameter rides a per-sample linear
ramp (20 ms default);
m_ramps_activemakes idle one integer test. The kernel test bounds the worst sample jump through a 100 ms preset recall and a per-sample 500→6000 Hz sweep — click-free is asserted, not assumed. - Preset morph in the kernel. 16 slots;
recall_presetisramp_toon all six parameters with a shared duration — as safe as any motion. - Anti-denormal on the stage states (
anti_denormal, thetap.comb~1e-15 idiom), insidetpt— a ringing-out filter otherwise decays into denormals and multiplies its own CPU cost. - Newton is guarded, not trusted. Seed clamp, derivative fallback,
iteration cap:
solve_exactcannot NaN or hang, only degrade towardsolver_fast. - Allocation-free after
prepare(); setters are plain stores into ramp targets, safe from the message thread while audio runs.
Checkpoint
Four trapezoidal one-poles put −180° and gain 1/4 at the prewarped cutoff;
negative feedback makes k = 4 the oscillation threshold, which is why
resonance is calibrated in quarters of k and self-oscillation lands on
pitch (8009 Hz for 8 kHz, measured). The linear loop solves in closed
form; the saturating loop is predicted linearly and committed through the
tanh stages once — exact in the small-signal limit, backstopped by a
guarded Newton solver. asym shifts the tanh operating point, comp
pre-multiplies away the derived 1/(1+k) droop, the Xpander table is
binomial algebra over the taps, and the oversampling chain pays tanh's
alias bill with a measured 13.5 dB. Every number in the user-facing
chapter traces to a line in this file.
The master phase and its corrections: vco.h
The user-facing chapter promised an oscillator whose folded harmonics sit ~47 dB down, whose analog section is "exactly zero by default" with a seed that works like a serial number, and whose FM survives through zero. Each claim is a theorem about this file or a measurement of it. This appendix derives the corrections — polyBLEP, the leaky triangle, the sync patch — then the analog-character section, including the one place where honest analysis contradicted intuition and the tests were written to match.
One phase, many readings
There is a single accumulator, m_phase ∈ [0,1), advanced once per sample
in step:
f_eff = base_hz · 2^(cents/1200) + fm_hz (pitch is exponential,
dt = f_eff / m_sr, clamped to ±0.49 FM is linear, in Hz)
adt = max(|dt|, 1e-8)
cents collects detune, drift, jitter, the per-unit tolerance offset, and
track — everything musical multiplies; only FM adds. Every waveform is a
reading of the same phase: sine through sin, saw as 2p−1, pulse as a
comparison against pw, triangle as an integral. The shape morph
crossfades adjacent readings of one phase, so it can never produce a
discontinuity the phase itself doesn't have. The problem is entirely the
discontinuities.
The fan-out the chapter title promises: every waveform is a reading of the same φ.
The residual: deriving poly_blep
A naive saw jumps by −2 at the wrap; a step's spectrum falls at only 6 dB/oct, so its harmonics march past Nyquist and fold back inharmonic. The ideal fix is a band-limited step — the integral of a sinc. The polyBLEP observation: the band-limited step differs from the naive step only near the edge, so instead of storing sinc-integral tables (minBLEP), approximate the difference with a polynomial. Take the crudest kernel one sample wide per side — a unit-area triangle, b(τ) = 1 − |τ| on τ ∈ [−1, 1] with τ in samples from the discontinuity — integrate, subtract the step:
s_bl(τ) = (τ+1)²/2 τ ∈ [−1, 0]
s_bl(τ) = 1 − (1−τ)²/2 τ ∈ [0, 1]
r(τ) = s_bl − s_naive:
r(τ) = (τ+1)²/2 τ ∈ [−1, 0) (before the edge)
r(τ) = −(1−τ)²/2 τ ∈ [0, 1] (after the edge)
Scale by 2 — the saw's wrap step — and these are exactly the file's
poly_blep branches: just after the wrap (t < dt, normalized t/dt = τ),
t + t − t*t − 1 = −(1−τ)² = 2·r; just before it (t > 1.0 − dt,
normalized (t−1)/dt = τ ∈ [−1,0)), t*t + t + t + 1 = (τ+1)² = 2·r.
Three properties fall out of the derivation:
- Continuity at the window edges. r(−1) = r(1) = 0: the correction fades in and out without new discontinuities of its own.
- The midpoint property. r(0⁻) = +½, r(0⁺) = −½: the corrected edge passes through the middle of the jump, as a true band-limited step does.
- "±1 sample" precisely. A sample lands in a branch iff its phase is
within
dtof the wrap, and phase movesdtper sample — exactly the last sample before and the first after each edge are touched, ever; the scope trace still looks like a saw.
The triangle kernel approximates the sinc (its spectrum is sinc², not a brickwall), so suppression is finite and measurable: the notebook drives a 3951 Hz saw, whose 13th harmonic folds to 3364 Hz, and measures it at −26.7 dB naive versus −74.2 dB with polyBLEP — 47.5 dB of suppression. The notebook's sample-level zoom shows the mechanism: the corrected saw passes through +0.588 and −0.856 on its way down, where the naive saw jumps in a single step.
Saw, pulse, and the second BLEP
saw_at is the reading minus the residual: 2·bent(p,·) − 1 − poly_blep(p, dt) (subtracted: the wrap step is −2, poly_blep is
normalized for +2). pulse_at is ±1 with two edges: the rising edge at
the wrap (step +2, residual added) and the falling edge at p = pw
(step −2, residual subtracted) — the latter evaluated at wrap01(p − pw),
re-centering the phase coordinate so that edge sits at zero of its own
window and the same branches apply. Calibration is pinned by measurement:
a bipolar pulse at duty d must average 2d − 1, and the notebook measures
−0.800 / −0.500 / +0.000 at 10 / 25 / 50 %; the kernel test holds the 25 %
mean within ±0.03.
The triangle: integrate, but leak
A triangle is the integral of a square — the classic analog trick. A ±1
square at frequency f forces slope ±4f (2 units in half a period 1/(2f)),
so the per-sample increment is ±4·f/fs = ±4·dt, which is tri_tick's
scaling exactly:
m_tri_state = 0.999 · m_tri_state + 4.0 · adt · sq
giving peak ±1 with no post-normalization. The 0.999 is the honesty tax:
the BLEP-corrected square's samples do not sum to exactly zero per period
(the two edges land at different sub-sample positions, so their
corrections don't cancel — and tri_pw skew under imperfect makes the
imbalance deliberate), and a pure integrator would ramp that residue to
infinity. The leak turns the integrator into a one-pole highpass with
corner fs·(1 − 0.999)/2π ≈ 7.6 Hz at 48 kHz — far below any audible
fundamental, high enough to hold DC bounded.
Why the integrated square is correctly antialiased: integration multiplies the spectrum by 1/ω, −6 dB/oct. The square's already-suppressed alias residual was generated near Nyquist, where 1/ω is smallest, so integrating a BLEP square improves the alias-to-harmonic ratio — the correction gets cheaper exactly where the waveform gets harder, which is why this hardware trick survives digitally intact.
Through-zero FM
Because FM adds in Hz after the exponential pitch math, dt can go
negative and the phase genuinely runs backward — that is all "through
zero" means, and why the sidebands stay coherent when the modulation
swings past the carrier. Two guards make it safe. The BLEP windows use
adt = |dt|: a window is a duration, one sample each side of an edge,
whichever way the phase travels (with a 1e-8 floor so the t /= dt
normalization survives a frozen phase). And dt is clamped to ±0.49: at
|dt| ≥ 0.5 the window tests t < dt and t > 1 − dt would overlap and
every sample would be "at an edge" — the clamp keeps the effective
frequency below Nyquist, where the model means anything at all. Measured:
a 500 Hz sine under ±900 Hz of FM at a 100 Hz rate — depth past the
carrier, genuinely through zero — puts its sideband at −13.7 dB with
−158.3 dB between the lines (144.5 dB of contrast), bounded at |y| = 1.00.
Hard sync, one-sided
A rising zero crossing on the sync input (m_sync_prev ≤ 0, sync > 0)
resets the phase. Linear interpolation locates the crossing inside the
sample:
frac = m_sync_prev / (m_sync_prev − sync) ∈ [0, 1)
m_phase = wrap01((1.0 − frac) · dt)
— the phase restarts from zero at the crossing and accumulates only the
remaining fraction of the sample, so sync pitch is sub-sample accurate
(the notebook's synced slave measures periodic at 110.1 Hz against a
110 Hz master). The reset is still a discontinuity of size
d = waveform_out_peek(p_old, …) − waveform_out_peek(wrap01(p_new), …),
sized on the morphed waveform without advancing the triangle integrator:
x = 1.0 − frac correction += d · 0.5 · x²
This is a first-order polynomial BLEP, honestly cruder than the saw's:
one-sided, because the pre-reset sample is already output when the
edge arrives — a reset cannot be predicted — and a one-sided patch can
never reproduce the full band-limited edge (the midpoint property needed
both sides). d·½x² is the triangle-kernel residual for a step landing x
into a sample; the code feeds it x = 1 − frac, the elapsed fraction
since the crossing, so its weighting runs opposite to the two-point post
branch (½·frac²) — at the first-order accuracy a one-sided patch can
claim, both are O(d) click reducers vanishing at one end of the window.
The header flags minBLEP tables as the wholesale upgrade; a m_pending
slot, read and cleared each sample but never written, is scaffolding for
the second correction sample it would need.
The analog section, derived (2026-07)
Two time scales of pitch noise. tick_drift is sample-and-hold noise
redrawn every m_sr / 2 samples (~2 Hz) smoothed by a one-pole with
a = 1 − exp(−2π·0.5/m_sr) — the exact discrete step of a 0.5 Hz lowpass.
tick_jitter is the same structure at ~80 Hz through ~40 Hz: the fast
companion, trembling where drift strolls. Both are depth-scaled in cents
into the pitch path. Measured: relative period spread 2.01×10⁻⁷ at jitter
0 versus 2.74×10⁻³ at 10 cents — four orders of magnitude of
micro-instability, still under the test's 0.02 ceiling.
The bent ramp, honestly. imperfect bows the saw via
bent(p, bend) = p + bend·p·(1−p), bend = 0.35·imp·m_tol_curve. The
parabola vanishes at both endpoints, so the wrap step stays exactly 2 and
the BLEP stays correctly sized — why the bend lives inside the ramp
reading. Now the honest part. In x = p − ½ the saw 2p−1 = 2x is odd (a
sine series) while the parabola p(1−p) = ¼ − x² is even (a cosine series):
the bend's Fourier content is in quadrature with the saw's own
components. Harmonic k gains an orthogonal part of relative size bend/(πk)
that moves its magnitude only at second order — about 0.05 dB at k = 1
for the maximal bend of 0.35, per √(1 + (bend/πk)²): a scope-obvious shark
fin, almost no harmonic-magnitude shift. Discovered by measurement — and
the kernel test matches the truth: it asserts the waveform bow (interior
deviation > 0.03 at imperfect 1) and makes no harmonic-magnitude claim.
Where the spectral work is actually done. The reset corner rounds
through a one-pole whose cutoff closes from ~22 kHz toward ~8 kHz
(fc = min(22000 − 14000·imp, 0.45·m_sr), coefficient cached against
m_round_imp): measured, the saw's 40th harmonic (17.6 kHz) is 6.4 dB
quieter at imperfect 1; the test requires > 4 dB. The triangle skews via
tri_pw = clamp(0.5 + imp·m_tol_tri·0.01, 0.05, 0.95) — duty asymmetry in
the integrated square is even harmonics — measured: triangle h2 rises from
−185.5 dB (numerically absent) to −34.1 dB at imperfect 0.8. The sine
reads a mildly bent phase (bent(p, 0.5·bend)), the pulse width takes a
static offset up to ±1.5 %, the whole unit a pitch offset up to ±2 cents
(m_tol_cents).
Which unit you own. The tolerances come from a separate stream:
compute_tolerances hashes the seed (m_seed * 2654435761u + 12345u)
into its own local LCG, never touching the runtime m_rng. The contract:
clear() resets m_rng = m_seed and all noise state but does not re-roll
tolerances — resetting the oscillator must never change which unit off the
production line you own; only set_seed re-rolls, because changing the
serial number is changing the unit. Every tolerance is scaled by
imperfect at use, yielding the contract the section rests on: at
imperfect 0, every seed is bit-identical to the ideal oscillator. The
kernel test renders seeds 7 and 8 and requires ya == yb — exact equality
over 24000 samples — and the notebook confirms it; conversely, with
drift 20 seeds 7 and 8 diverge by up to 0.183 (measured) while the same
seed renders bit-identically, and the test pins that at imperfect 0.6
different seeds are audibly different units.
track.
cents += track · log₂(base_hz / 440.0)
A V/oct converter's calibration error grows linearly in octaves from its
trim point; this is that line, exact at 440 Hz by construction
(log₂ 1 = 0). Measured at track 5: −15.0 / −10.0 / −5.0 / +0.0 / +5.0 /
+10.0 / +15.0 cents across −3…+3 octaves; the test holds the trim point
under 1 cent, ±3 octaves within 2 cents of ±15.
The engineering ledger
- Determinism is structural. All randomness flows from one 32-bit LCG
(1664525 / 1013904223) seeded by
m_seed(0 remapped to 1); no wall clock, nostd::random— renders, tests, and mc. stacks reproduce bit-for-bit, and the jitter test pins same-seed bit-identity. - Off means exactly off.
tick_drift/tick_jitterreturn before consuming the RNG at zero depth — a default-configured oscillator never advancesm_rng, so different seeds render identically until a stochastic feature is engaged. The corner-rounding pole is gated onimp > 0.0, its state primed while bypassed (m_round_lp = y) so engagingimperfectmid-note is click-free. - The triangle integrator ticks only when the morph needs it.
waveform_outshort-circuits the crossfade endpoints (a <= 0.0), andtri_tickis stateful — skipping it when unused is a cost saving and a correctness rule (parking at pure saw must not silently integrate); sync sizing useswaveform_out_peekfor the same reason. - Clamps with reasons:
dtat ±0.49 (window overlap / Nyquist),adtfloored at 1e-8 (division inpoly_blep),tri_pwin [0.05, 0.95] andpwin [0.01, 0.99] (an edge pair must stay two distinct windows). - The house frame: per-sample linear ramps with an active-count fast
path, 16 preset slots morphable over time, allocation-free processing,
setters safe while audio runs — the same bones as
ladder.handsvf.h, so the wrapper stays a shim.
Checkpoint
One master phase; every waveform is a reading of it, and every reading's
discontinuity gets the residual of a triangle-kernel band-limited step —
two samples per edge, continuous at its window boundaries, a measured
47.5 dB of alias suppression at the folded 13th harmonic. The triangle
integrates the corrected square (slope ±4f, hence 4·adt) with a 0.999
leak; FM adds in Hz so the phase can run backward, adt keeping the
windows directionless; sync resets with sub-sample accuracy and patches
the step one-sidedly because resets can't be predicted — minBLEP is the
flagged upgrade. The analog section is derived noise at two time scales, a
quadrature-honest bent ramp, a rounding pole, a skewed duty, and a
calibration line exact at A440 — all drawn from a tolerance stream
clear() never touches, all scaled by imperfect, and all provably
absent at zero: the ideal oscillator is a test-pinned invariant, not a
default setting.
Detector, law, and a borrowed filter: autowah.h
The user-facing chapter (The pedal that listens) made three
measurable claims: the sweep law matches its design to 0.000 cents, the
follower's timing is an honest RC discharge, and sensitivity at the floor
turns the object into a truly fixed filter. This appendix derives each,
starting from the decision that tap.autowah~ is mostly not a new filter.
The composition decision: don't write a second SVF
The Snow White's core is an LM13700 OTA state-variable filter swept by an
envelope. TapTools already ships an SVF kernel whose defining property —
proved in the SVF appendix — is A-stability under per-sample
cutoff modulation. An envelope-swept filter moves its cutoff every sample by
construction; the property the wah needs most is exactly the one svf.h
already guarantees by theorem. So wah_filter owns a
tap::tools::svf::svf_filter member (m_svf) and drives it through the
signal-rate path, m_svf.tick(m_cutoff) then m_svf.process(0, x), once
per sample. The house rule makes this legal: objects under
source/projects/ stay self-contained, but inside the kernel repo sharing
between kernels is encouraged — autowah.h simply #include "svf.h".
Composition buys something subtler than saved code: the corner-identity
argument. When the envelope is off, the wah is a bare SVF, and that is a
testable equation rather than a resemblance. The kernel test
("sensitivity at the floor is the cocked-wah: bit-close to a bare svf at
bias") runs the wah at sensitivity −60 dB, bias 800 Hz, resonance 0.7
against a separately constructed svf::svf_filter fed
ref.process(x, 800.0), and requires maxerr < 1e-12 over half a second of
signal. The wet paths are arithmetic-identical — same tick(cutoff) entry,
same clamps, same solve — so why 10⁻¹² and not ==? The dry leak: at
mix = 100 the mix angle is θ = π/2, and while sin(π/2) is exactly 1.0 in
doubles, cos(π/2) rounds to 6.123×10⁻¹⁷, so the output carries a
6×10⁻¹⁷·x dry residue. The tolerance covers one ulp-scale cosine, nothing
else. Resonance meaning is shared the same way — the wah's 0..1 knob goes
through the SVF's own q_from_resonance mapping, so "resonance 0.7" means
the same Q in both objects.
The composition decision, drawn: everything amber is this file; the blue box is svf.h, borrowed intact.
The detector: gain, rectifier, follower
The detector chain in process() is three lines:
driven = key · m_sens_gain (dB → linear input gain)
rect = |driven| (full-wave, default)
or max(driven, 0) (half-wave, the traced single-diode topology)
m_env += coef · (rect − m_env) (one-pole follower)
The sensitivity floor is a contract, not a clamp. update_derived maps
the dB knob as
m_sens_gain = 0 if sens_db ≤ −60 dB
= 10^(sens_db / 20) otherwise
−60 dB is not "very quiet" (that would be gain 0.001); it is exactly zero.
With m_sens_gain = 0 the rectifier output is identically 0, m_env decays
to 0 and stays there, m_sweep = tanh(0) = 0, and map_cutoff(0) returns
m_bias exactly — the test asserts w.cutoff_hz() == 250.0 with ==. That
is what makes the pedal's secondary "cocked wah" mode a true fixed filter
rather than an approximately-fixed one that still breathes a few cents with
the input. Factory slot 3 is that voicing as data.
The follower coefficient is the exact RC discretization. The analog detector is a capacitor charged toward the rectified signal: env′ = (rect − env)/τ. Solving that ODE exactly over one sample period T = 1/fs gives env[n] = rect + (env[n−1] − rect)·e^(−T/τ), which rearranges to the code's recurrence with
a = 1 − e^(−1/(τ·fs)) (the file: m_attack_coef = 1 − exp(−1000/(ms·m_sr)))
After N = τ·fs samples of a step, the remaining error is
(1−a)^N = e^(−N/(τ·fs)) = e^(−1): the envelope reaches 63.2% in exactly τ.
This is not the cheap approximation a ≈ 1/(τ·fs); the exponential form makes
the ms parameters honest at any rate. The notebook measured it: attack set
to 2.0 ms reaches 63% in 1.94 ms; decay set to 250 ms falls to 36.8% in
256 ms; and a log-domain fit of the release is a pure exponential with
τ = 252 ms and residual σ = 0.004 — an RC discharge, like the hardware.
The attack/release asymmetry is one branch, coef = (rect > m_env) ? m_attack_coef : m_decay_coef — the diode charges the cap through one
resistance and lets it bleed through another.
Full-wave default, half-wave option. The traced hardware detector is a
single diode: it charges only on positive half-cycles, so the envelope
droops between charges at the signal's fundamental. Full-wave rectification
charges twice per cycle and halves the gaps. The follower cannot filter this
out without also slowing the response, so the ripple rides the envelope and
frequency-modulates the cutoff — the hardware's "sweep-rate ripple." The
notebook quantified the A/B on a 110 Hz tone (decay 60 ms): settled envelope
ripple (std/mean) 0.7% full-wave vs 2.6% half-wave. The kernel defaults
to the cleaner full-wave and keeps half-wave selectable (set_rectifier())
because the flavor question is a hardware-listening question, not a math
question — it waits for the calibration pass.
The sweep law, and where it is honest about ignorance
m_sweep = tanh(k_env_knee · m_env) k_env_knee = 1.5
m_cutoff = m_bias · 2^(m_sweep · m_range) clamped to [20 Hz, min(20 kHz, 0.45·fs)]
Three deliberate choices:
(a) Exponential in Hz. Equal envelope increments move the cutoff by
equal octaves, which is how a sweep sounds uniform — pitch perception is
log-frequency. The honest caveat lives in the header and in
map_cutoff()'s own comment: the LM13700's frequency is linear in control
current, so the pedal's true law hinges on the BJT stage that converts the
envelope voltage into that current — plausibly exponential (a BJT's
collector current is exponential in V_BE), but not yet measured. That is why
the law is one isolated function: if calibration finds a linear V→I driver,
map_cutoff becomes bias + sweep · span_hz and nothing else in the kernel
changes.
(b) The tanh soft knee. Without it, hard playing would pin sweep at a
clamp rail — a hard corner in the control trajectory, audible as the filter
slamming its ceiling. With it, the ceiling is approached asymptotically.
The arithmetic at the defaults (bias 250 Hz, range 3.3 octaves,
sensitivity 0 dB):
full-scale DC key → m_env → 1
m_sweep → tanh(1.5) = 0.905
m_cutoff → 250 · 2^(0.905 · 3.3) ≈ 1982 Hz
— about 2 kHz, under the asymptotic rail 250·2^3.3 ≈ 2462 Hz, which itself
matches the hardware's published 250–2500 Hz span. The unit test pins the
settled cutoff into (1800, 2100) Hz and separately drives an absurd +24 dB
sensitivity into an 8× full-scale key to confirm the cutoff saturates at the
ceiling (reaching > 99% of it) instead of running away. range is signed —
negative sweeps down from bias, a deliberate extension the pedal never had.
(c) Measured. The validation notebook swept the envelope range and compared measured cutoff against the designed curve: max error 0.000 cents. The law in the code is the law on paper; when hardware recordings arrive, any disagreement is a fact about the model choice, not about the implementation.
One filter, two owners: the forwarding discipline
Both kernels ship per-sample parameter ramps. Run both and every set would
be smoothed twice — lagged, and worse, shaped (a ramp of a ramp is not a
ramp). So prepare() declares a single owner:
m_svf.set_smooth_ms(0.0); // this kernel owns all smoothing; svf setters snap
The wah's own ramp array smooths every audible parameter; the composed SVF's
setters snap instantly to whatever the wah forwards. And forwarding is
change-gated: update_derived pushes set_resonance / set_drive /
set_circuit only when the cached values (m_svf_resonance, m_svf_drive,
m_svf_circuit, seeded to −1 to force the first forward) actually differ.
The reason is the SVF's two-tier update from its own appendix: those setters
dirty the shape tier (damping, mix weights, drive gain — a pow and the
mix logic). While any wah ramp is active, update_derived runs every
sample; forwarding unconditionally would re-run the SVF's shape update every
sample of every bias morph even though resonance never moved. Cutoff needs
no gate at all — tick(cutoff_hz) lands in the SVF's cutoff cache, which
recomputes the tan and solve constants only when the value differs.
The circuit switch
drive at 0 dB runs the SVF's clean linear circuit; anything above engages
circuit_driven — tanh band-node limiting, 2× oversampled — as the optional
OTA-flavored color stage. The switch is a threshold in update_derived:
circuit = (m_svf_drive > 1e-6) ? circuit_driven : circuit_clean
Why is switching circuits mid-stream acceptable? At the switch point drive ≈ 0 dB, so the driven circuit's input gain is 1 and tanh is near-identity at typical band-node levels — the two circuits compute nearly the same output, and the transition is benign. Honestly stated: near-identity, not identity. At high resonance the band node runs hot, tanh visibly bends, and the driven circuit also brings its oversampling path with it — so engaging drive from zero on a screaming resonant setting is a small audible step. The abuse test accepts this trade explicitly: resonance 1.0 plus max drive on square-wave bursts must stay bounded and finite (the SVF's bounded self-oscillation doing its job), not polite.
Output staging is the equal-power crossfade shared with tap.crossfade~:
θ = mix·π/200, m_dry_gain = cos θ · g, m_wet_gain = sin θ · g — the
master gain g rides both paths, so gain never changes the balance.
The engineering ledger
- Per-sample cost accounting. Settled steady state pays: one rectify,
one follower multiply-add, one
tanh(knee), oneexp2(law), the SVF solve, two mix multiplies. Thepow(10, ·)/expcalls live only inupdate_derived, which runs per sample while ramps move and exactly once after they settle (m_derived_dirtyre-arms only whenm_ramps_active > 0). The real recurring cost is the SVF's cutoff-tiertan, paid whenever the envelope actually moved the cutoff — and skipped by the SVF's cutoff cache whenever it didn't (silence, or the cocked wah). envelope()andcutoff_hz()exist for measurement. They readm_sweepandm_cutoffafter the fact; the C ABI'staptools_wah_process(..., env_out, cutoff_out, n)taps them per sample, and the validation notebook'strace=Truepath — the ground truth its STFT peak-trajectory extractor was proven against (0.979 log-frequency correlation) — is built on exactly these two accessors.- The preset-morph engine is the GRM pattern (16 slots, the
grm_comb/grm_pitchaccumhouse count):store_presetcaptures ramp targets (knob positions, never mid-ramp instantaneous values),recall_preset(slot, seconds)re-targets every ramp so a morph is just nine simultaneous ramps — re-targeting mid-morph stays continuous for free. The test walks a 100 ms morph and requires bias to move monotonically with no step larger than one ramp increment. - Factory slots are data, not code: four
paramsstructs (guitar / bass / slow swell / cocked wah) in slots 0–3. Changing a voicing after the hardware session edits numbers, not logic; the test pins them. - Structural switches (
mode,rectifier) are not ramped or morphed — interpolating between rectifier topologies has no physical meaning. - Anti-denormal on the envelope (
< 1e-15 → 0, thetap.comb~idiom): a decaying exponential otherwise glides into denormals and multiplies its CPU cost during silence. - Sidechain by signature:
process(x)isprocess(x, x); the wrapper's key inlet is the two-argument form. Single-channel by design — per-channel envelopes are the correct behavior undermc.wrapping.
The calibration pass, by construction
Every open hardware question maps to one isolated switch point: the sweep
law is map_cutoff() (one function), the stock filter tap is m_mode's
default (mode_lowpass, flagged in the header as inference), the detector
topology is set_rectifier(), the knee is k_env_knee (one constant). The
validation notebook is the instrument that will close them: its extractor
recovers the swept peak from wet audio alone, is already calibrated against
the kernel's own trajectories, and its last cell waits for
snowwhite_*.wav. When the pedal arrives, disagreements land on named
constants — not on a rewrite.
Checkpoint
The wah is a detector and a law in front of a borrowed filter. Composing
svf_filter puts the per-sample-modulation stability where it is already
proven, and makes "sensitivity off equals a bare SVF" an identity checked to
10⁻¹² (the gap being one rounded cosine). The follower coefficient
1 − e^(−1/(τ·fs)) is the exact RC step — 63.2% in τ by algebra, 1.94 ms
measured for 2.0 set. The sweep law is exponential-in-Hz through a tanh knee
that turns hard playing into asymptotic approach (~2 kHz at the defaults,
under the hardware's 2.5 kHz rail) — measured at 0.000 cents against design,
and honestly provisional, isolated in one function until the real pedal
votes.
Convolution without compromise: conv_engine.h
The user-facing chapter (Borrowed rooms) made a flat claim:
tap.convolve~ is exact — not "high quality," exact — and its only cost
is a latency of precisely blocksize samples. Claims like that are either
algebra or advertising. This appendix does the algebra: why the convolution
is partitioned at all, why overlap-save, why the latency is exactly one
partition, and how an impulse response can be replaced mid-performance
without the audio thread ever seeing a torn table.
Why partitioned: the cost triangle
Direct convolution of a stereo pair against an L-second IR at rate fs costs
cost_direct = L·fs MACs per output sample per path
— 48 000 multiply-adds per sample for one second of room at 48 kHz, times four paths for true stereo. Untenable. The classical fix is to convolve in the frequency domain: transform the whole IR once, multiply spectra, invert. But a single-FFT scheme cannot emit anything until it holds a full frame of input, so its latency equals the IR length — seconds of delay for a reverb. Also untenable.
Uniform partitioning takes the middle of the triangle. Split the IR into P
equal partitions of B samples, h_j[k] = h[j·B + k]; then by linearity
h = Σ_j h_j delayed by j·B
Each partition is short enough that its convolution can be computed with a
small FFT once per B-sample block, and the delays j·B are whole blocks —
which, we will see, cost nothing but indexing. FFT economics, latency of one
partition. This is the standard engine of the genre for a reason.
Overlap-save: the framing, derived
The engine convolves each partition by circular convolution over an FFT of
size m_fftsize = 2·m_block — size 2B for partition size B. Circular and
linear convolution are not the same thing; the design question is which
output samples of the circular product are also the linear ones.
Take the frame the code actually builds in process_block():
frame_m = [ block_{m−1} ; block_m ] (m_fre[j] = m_prev[ch][j]; m_fre[B+j] = m_inblk[ch][j])
and a partition h zero-padded from B to 2B. The circular convolution is
y_circ[n] = Σ_{k=0}^{B−1} h[k] · frame[(n − k) mod 2B]
For n in the second half, n ∈ [B, 2B), and k < B, the index n − k stays
in [1, 2B): the mod never wraps, so y_circ[n] equals the linear
convolution of h with the input stream at that time. For n < B the mod does
wrap, splicing in samples from the frame's far end — time aliasing. So each
2B-point product yields exactly B valid samples, the second half, and the
code keeps precisely those:
m_outblk[oc][j] = m_are[m_block + j]; // overlap-save: discard the aliased first half
That is the whole scheme: hop by B, keep the clean half, discard the dirty
half. The alternative, overlap-add, zero-pads each input block instead and
sums overlapping output tails — equally exact in theory, but it carries a
partial-sum accumulation buffer across block boundaries, one more piece of
state to get right. Overlap-save's output block is finished the moment the
IFFT returns: no summation state, no output windowing, no crossfading —
there is nothing between the IFFT and the output buffer that could be
inexact. Every step in the chain — framing (a copy), FFT and IFFT (the
shared radix-2 in fft.h, whose inverse divides by N so a round trip
reconstructs its input), and the multiply-accumulate — is exact linear
algebra in double precision. The engine has no tuning parameters that trade
accuracy for speed; its error budget is rounding noise, and the measurements
below confirm that is all there is.
The frequency-domain delay line
Why FFT cost is constant in IR length: the transforms bracket the structure, and only the MAC sees the partitions.
The delays j·B remain. Delaying partition j's contribution by j blocks is
the same as convolving it with the input from j blocks ago — and the frame
for block m − j has already been transformed. So the engine keeps a ring of
past input spectra (the FDL, m_fdl_re/m_fdl_im, one per input channel)
and forms the output spectrum as
Y_m[k] = Σ_{j=0}^{P−1} H_j[k] · X_{m−j}[k]
which is exactly the inner loop: slot = cur − p (mod m_max_parts), then
a complex multiply-accumulate over all m_fftsize bins into
m_are/m_aim. The consequence for cost is the design's payoff: one
forward FFT per input channel and one inverse FFT per output channel per
block, regardless of P. Growing the IR grows only the MAC. Per output
sample, the MAC costs P·2B complex MACs / B samples = 2P complex MACs ≈ 8P
real multiplies per path, versus P·B real MACs for direct convolution — a
factor of B/8 (64× at B = 512), with the FFTs an O(log B) constant on top.
Latency is exactly B, by the framing
Follow one sample through process(): it is written into m_inblk at
position m_pos, and the output handed back at that same call is read
from m_outblk[m_pos] — a block computed when the previous block
completed. When block m finishes, process_block() runs with a frame ending
at the newest sample, and its B valid outputs are the linear convolution up
to that sample; they are then dealt out during block m + 1. So output sample
t carries y(t − B): the first partition's contribution to a block is
computed from the block just gathered, not from anything older, and the
delay is one partition — no more (the frame includes the newest sample) and
no less (nothing can be emitted before a block is complete). The unit test
pins both edges: the first B output samples are silence (the pre-roll), and
an IR of δ at index 5 yields the input delayed by exactly B + 5. The
notebook verified the latency at B = 64, 256, and 1024 — always exactly B.
Exact, measured
The notebook (executing the real engine through the C ABI) puts numbers on "exact": against a direct time-domain convolution of the same IR — 24 000 samples, B = 512, 47 partitions — the maximum difference is 3.13×10⁻¹². Across block sizes the outputs agree with the direct reference to 2.27×10⁻¹³ / 1.17×10⁻¹² / 1.76×10⁻¹² (B = 64/256/1024), and with each other, latency-removed, to 1.81×10⁻¹² — the block size is a CPU/latency dial with no audible existence. An impulse through the engine reconstructs the loaded IR to 5.5×10⁻¹⁴, and a synthetic 0.60 s-RT60 reverb measures back at 0.599 s. The unit test does the same job in CI with independent per-path IRs, deliberately awkward 10-sample process chunks that straddle block boundaries, and 10⁻⁹ tolerances.
True stereo: four paths, two FFTs
A stereo room is a 2×2 linear system, and the engine runs all of it:
out_l = in_l ∗ h_LL + in_r ∗ h_RL ; out_r = in_l ∗ h_LR + in_r ∗ h_RR
with path = in_channel·2 + out_channel (0 = LL, 1 = LR, 2 = RL, 3 = RR).
The economics are better than 4× mono: the two input FFTs are shared across
all four paths, and each output channel needs one inverse — so a block costs
2 forward FFTs, 2 inverse FFTs, and 4 MAC passes. The notebook pins the
routing: an impulse into L only emerges on R at exactly the cross-feed
path's gain (0.600 expected, 0.600 measured) and, with the off-diagonal
paths silent, R stays at 0.0 — cross-terms cannot hide in each other. The
mapping from buffer~ channel count to paths (4+ = true stereo, 2 = dual
mono, 1 = same room both sides) is wrapper policy; the engine only ever
knows four pointers, any of which may be null for a silent path.
The atomic IR swap
Loading a room while the music plays is the one place this engine touches
concurrency, and it is confined to a single atomic. The IR tables are
double-buffered per path (m_ir_re[path][slot], slots 0/1). load_ir() —
which runs off the audio thread; it is the expensive part, P analysis FFTs
per path — writes only the inactive slot, then publishes:
m_slot_parts[inactive] = P; // written before the publish...
m_active.store(inactive, std::memory_order_release); // ...so (slot, P) stay consistent.
The perform loop does one acquire load of m_active per block. The
release/acquire pair means that if the audio thread observes the new slot
index, it also observes that slot's fully written spectra and its
partition count — slot and P travel through one atomic, so there is no
window where the loop MACs over half-written tables or over the wrong
number of partitions. Until the store, the loop reads the old slot, which
the loader never touches. The discipline is single-writer double-buffering:
publishes are serialized through the wrapper's message path, and the just
vacated slot is only rewritten by the next load.
Why does a swap settle in exactly one block? Because the FDL stores input
spectra, not output. The first process_block() after the publish already
renders the entire tail — all P partitions — from the new IR against the
existing input history; the only samples that differ from a
new-IR-from-the-start engine are the ones sitting in m_outblk, computed
just before the swap. The unit test pins the settling (output equals the new
IR's pure delay from shortly after the swap); the notebook pins it exactly:
max |swapped − reference| after +1 block = 0.00×10⁰ — bit-identical to
an engine that had the new IR from the start — with RMS continuity across
the swap instant, 21.922 before, 22.147 just after, no dropout. Honest
limit: the swap is a hard splice between two exact convolutions, click-free
but discrete; the engine does not interpolate between rooms.
A related freebie of the input-side FDL: analysis of incoming audio happens before the has-IR check, so the delay line is warm even while no IR is loaded — load the first room mid-stream and its tail renders immediately from audio already played.
The engineering ledger
- Uniform, not Gardner non-uniform, partitioning. Non-uniform schemes (short partitions first, growing later) can push latency below B for the same CPU, but they need multiple FFT sizes and a scheduler that spreads long-partition work across blocks — real complexity with real failure modes. The object's latency budget is satisfied by making B small (64 samples = 1.3 ms, verified exact above); complexity was not bought that nothing needed.
- The shared FFT.
fft.his the in-house radix-2 Cooley–Tukey used by the whole spectral set; per its header it lived byte-identical insideconv_engine,tap.nr~, andtap.spectra~before being consolidated at the kernel split. In-place, forward unscaled, inverse divides by N — round-trip exact by construction. No external FFT dependency, per the porting philosophy. - IR stored as float32, deliberately. A
buffer~holds 32-bit samples; the engine quantizes atload_ir(static_cast<double>(src[idx]) * scale) and computes in doubles thereafter. The notebook's direct-convolution reference casts its IR through float32 the same way — so the 10⁻¹² figures isolate the algorithm, not the source quantization the wrapper inherits from Max regardless. - Geometry only in
configure(). Partition size and capacity determine every buffer, so reallocation happens only where the audio thread is idle — the wrapper calls it fromdspsetup.clear()flushes running state withstd::fillonly, no reallocation, and is safe from a message handler;process()allocates nothing, ever (scratch spectram_fre/m_fim/m_are/m_aimare preallocated and reused). - Capacity vs. length. The FDL ring is sized
m_max_parts; a loaded IR usesP ≤ m_max_partspartitions (load_irclamps), and the MAC runs over P only — a short room in a big engine costs a short room. - The deferred optimization, on the record. The spectra are stored and MAC'd full-complex; the input is real, so a half-spectrum (N/2 + 1 bins, Hermitian symmetry) form would halve both the MAC and the IR/FDL memory. The header flags it and parks it, under the same house rule the SVF appendix recorded: optimizations land bit-identical or explicitly signed off — and re-deriving the packing arithmetic is exactly the kind of change that gets signed off with a measurement, not slipped in.
Checkpoint
Partitioning splits the IR by linearity; overlap-save framing makes each 2B-point circular product yield B exactly-linear samples with nothing to window or crossfade; the frequency-domain delay line turns partition delays into ring indexing, so FFT cost is constant in IR length and only the MAC grows. Latency is one partition by the framing — measured at exactly B for every B tried — and exactness is measured at 10⁻¹²-and-below everywhere it can be probed. The one concurrent act, swapping rooms, rides a single release/acquire atomic over double-buffered tables, and settles bit-identically in one block because the delay line remembers input, not output. The compromises the genre usually accepts — approximate tails, block-size coloration, swap dropouts — are absent, and the measurements say so.
Ring time as the truth: grm_comb.h
The user-facing chapter made three claims that sound like
marketing until you do the math: that a voice keeps its decay as its pitch
sweeps, that warp stretches the partials while the fundamental stays in
tune, and that phase at 100 cancels the even harmonics — exactly. This
appendix derives all three the way the file was designed, then walks the
code-level decisions. The behavioral claims below are pinned by the kernel
scenarios in tap.5comb_tilde_test.cpp, which drive
tap::tools::fivecomb::comb_bank directly (no Max in the loop); the few
numbers outside the test suite are marked as measured on the kernel for
this chapter.
A comb is a string
One voice is a delay of d samples fed back on itself. Ignoring the in-loop
filters for a moment, the recursion the code implements
(y = in + fb * ap_out, written back into the delay line) is:
y[n] = x[n] + fb · y[n − d] h[n] = δ[n] + fb·δ[n−d] + fb²·δ[n−2d] + …
The impulse response is echoes every d samples, each scaled by another factor of fb — and echoes every 1/f seconds is a tone at f and its harmonics: a plucked string tuned to f = fs/d. The first kernel scenario pins the geometry: a 500 Hz voice at 48 kHz (d ≈ 96) puts its first three echoes at one period spacing, within a couple of samples of 96/192/288.
Ring time as the truth
The legacy object exposed fb directly. This file refuses to, and the reason
is in the impulse response above. After t seconds the signal has made t·fs/d
round trips, so the decay envelope is level(t) = fb^(t·fs/d).
Reverberation's standard yardstick is RT60, the time to fall 60 dB — a factor of 10^(−60/20) = 10⁻³. Set level(rt60) = 10⁻³ and solve:
fb^(rt60 · fs / d) = 10⁻³ ⇒ fb = 10^(−3·d / (rt60·fs))
which is character for character the file's line in update_derived():
m_fb[v] = min(pow(10.0, −3.0·d_total/(rt60·m_sr)), k_fb_max).
The res knob maps to rt60 on a log curve before this — rt60 = k_rt60_min · (k_rt60_max/k_rt60_min)^(res_eff/100), 20 ms at res → 0⁺ up to
100 s at res 100 — because equal knob travel should mean equal ratios of
decay time, which a linear map to fb spectacularly is not (all the action of
a raw-feedback comb lives in the last few percent of the knob).
Now the design consequence, which is the chapter title. Because fb is
re-derived from the current delay every time update_derived() runs,
the ring time is the invariant: sweep a voice's frequency and fb is silently
re-solved to hold rt60 constant. A raw-feedback comb has it backwards — hold
fb fixed and rt60 = −3d/(fs·log₁₀ fb) is proportional to d, so low notes
ring 1/f longer and high notes choke. The "resonance maps to ring time"
scenario pins the calibration: inverting the log curve for rt60 = 1 s gives
res ≈ 45.93, and the measured tail drops close to the ideal 30 dB over a
half-second window (the test accepts 22–40 dB; it lands near 32, the excess
being upper partials that the interpolator and loop lowpass damp slightly
faster).
Why the delay must be fractional
At 48 kHz a 440 Hz comb needs d = 48000/440 = 109.09 samples. Round to 109 and the voice plays 48000/109 = 440.37 Hz — about +1.4 cents. Worse than the absolute error: each of the five voices quantizes differently, so the carefully-tuned beating between voices (the point of a bank) is replaced by whatever the rounding produced. The legacy abstraction had integer delays and control-rate stepping; the file header names that as the main reason it never sounded like the GRM original.
So the tap is fractional — but not linear. Reading between samples with linear interpolation is a two-tap filter H(z) = (1−η) + η·z⁻¹ where η is the fractional part, with magnitude:
|H(ω)|² = 1 − 2η(1−η)(1 − cos ω)
— a lowpass whose damping depends on η (worst at η = ½, where |H| = cos(ω/2),
a null at Nyquist). Two failure modes follow. Statically, this filter sits
inside the loop: its droop is applied once per round trip and compounds
into the ring, so two voices with different fractional parts get different
brightness decay for free. Dynamically, a sweep cycles η through 0 → 1
repeatedly, so the loop's damping ripples at the sweep rate — audible
dulling and level flutter. read_hermite() is the 4-point, 3rd-order
Hermite (Catmull-Rom) interpolator instead: C¹-continuous, passband flat to
far higher frequency, far weaker dependence on η. Its cost is the geometry
constraint noted in the code — the youngest of its four points is one ahead
of the base, so d must exceed 2 strictly, hence k_min_delay_samples = 2.5
and the ceiling f_ceil = min(k_freq_ceil_hz, m_sr / k_min_delay_samples).
The feedback chain, in order
Per sample, the loop path in comb_voice::process() is:
delayed = read_hermite(d_read) → one-pole lowpass → DC blocker
→ warp allpass → × fb → + in → write
The lowpass is the string's brightness decay: every round trip gets a
little darker, highs first, like a real string. Its coefficient is exact —
m_lp_a[v] = 1 − e^(−2π·fc/m_sr) — placing the −3 dB corner at fc by
construction (the one-step discretization of an RC section). The in-file
comment flags the deviation: tap.comb~ used hz·2/sr, which is not even
the small-argument limit of the exact map (that would be 2π·fc/fs) — its
actual corner lands near fc/π, a factor-of-three tuning error on a labeled
frequency knob. Faithful porting stops where the parameter lies about its
units.
The DC blocker (y = norm·(x − x1) + R·y1, R = k_dc_block_r = 0.999,
norm = k_dc_block_norm = (1+R)/2, ~7 Hz corner at 48 kHz) replaces the
legacy tap.comb~ hard ±1 autoclip — the file's most consequential
retirement. The clipper existed to stop runaway; but a clipper in a resonant
loop is a distortion stage, and at high resonance — precisely where the
GRM sound lives — the legacy object audibly distorted. The modern argument:
cap fb below unity (k_fb_max = 0.99999), kill the loop's DC transmission
(the blocker's zero at z = 1), and the linear loop contracts — no limiter
needed, so res 100 rings clean.
The norm factor is this chapter's own contribution, and the story is worth
a paragraph. The raw blocker (1 − z⁻¹)/(1 − R·z⁻¹) is not passive: its
magnitude peaks at 2/(1+R) ≈ 1.0005 toward Nyquist. While proving the
contraction claim for the first edition of this chapter, the measurement
came back false in one corner: with the loop lowpass wide open the product
|H_lp·H_dc| crossed unity near 450 Hz, and a voice tuned there at res 100 —
where fb saturates at the cap — measurably swelled at ~+0.2 dB per second.
The fix is the normalization: scaling the blocker by (1+R)/2 pins its peak
gain at exactly 1 (the zero at DC is untouched), so with the allpass at unit
magnitude and the one-pole lowpass ≤ 1, the loop gain is bounded by fb alone
and fb < 1 now really is the airtight inequality. A kernel test pins the
formerly-failing corner: 450 Hz, res 100, lp at 20 kHz, twelve seconds of
ring-out, decaying window over window.
The normalization has a side effect the file also pays for: the blocker now
slightly attenuates each voice's fundamental (a few parts in 10⁴ at mid
frequencies, more for very low voices), which — uncompensated — would shave
the top off long ring times. So update_derived divides the RT60-derived fb
by dc_block_gain(ω₀), the blocker's magnitude at the voice's fundamental —
the same pay-the-fundamental-back philosophy as the warp compensation below,
and clamped to k_fb_max so the contraction bound survives. The second new
kernel scenario pins the payoff: an impulse-excited voice at res 50 measures
its RT60 within 10 % of the map's 1.41 s target.
warp: dispersion, and paying the fundamental back
The allpass is the modern GRM Comb's character control. warp sets
m_ap_c = −k_warp_coef_max · warp/100 (c ∈ [−0.85, 0])
and inserts H(z) = (c + z⁻¹)/(1 + c·z⁻¹) into the loop — unit magnitude
everywhere (it cannot alter the decay), pure phase. At c = 0 it degenerates
to z⁻¹, an honest one-sample delay: warp 0 is exactly the harmonic Classic
comb. Its phase delay in samples is what allpass_phase_delay() computes:
τ(ω) = [ atan2(sin ω, c + cos ω) − atan2(c·sin ω, 1 + c·cos ω) ] / ω
τ(0) = (1 − c)/(1 + c) (the DC limit the w < 1e−9 branch returns)
For negative c, τ falls monotonically with frequency — at c = −0.85, from (1.85/0.15) ≈ 12.3 samples at DC down to 1 sample at Nyquist. A partial's resonant frequency is set by its total round-trip time d_read + τ(ω), so upper partials, seeing a shorter loop, land sharp of the harmonic series — the stretched partials of a stiff piano string, exactly the physics that motivates the control.
Left there, the fundamental would sharpen too. The compensation is one line:
m_d_read[v] = max(d_total − ap_tau, k_min_delay_samples), where
ap_tau = allpass_phase_delay(m_ap_c, 2π·f/m_sr) is evaluated at the
voice's fundamental. The main tap is shortened by precisely the phase
delay the allpass adds at that frequency, so the fundamental's round trip is
d_total again — pitch stays put while the overtones stretch. The max is
the documented physical limit: at extreme warp × high tuning, d_total − τ
falls below the interpolator's 2.5-sample floor, the loop cannot get shorter
than the dispersion, and the pitch flattens — physical, and flagged in the
maxref. The warp scenario pins the endpoints: warp 0 echoes at one period;
warp 100 stays bounded, still resonates, and its tail correlates < 0.5 with
the harmonic tail (a genuinely different spectrum, not a filter tilt).
phase: plucking the string at its midpoint
The output tap is out = y − pickup·read_linear(d_half) with
d_half = d_total/2. Consider loop content at harmonic n of the voice —
frequency n·f, i.e. n·fs/d. Delaying it by d/2 samples shifts its phase by
Δφ = 2π · (n·f/fs) · (d/2) = n·π
Even n: Δφ is a multiple of 2π, the delayed copy equals the original, and the subtraction (at pickup = 1) cancels it exactly. Odd n: Δφ = π, the copy is inverted, and subtraction doubles it. Pickup therefore sweeps continuously from the full series to odd-harmonics-only — plucking a string at its midpoint, where the even modes have a node. The test uses a 1 kHz voice at 48 kHz so the half tap lands on exactly 24 samples; Goertzel measures the 2f-to-f power ratio collapsing below 5% of its phase-0 value.
The parameter engine
All 22 parameters (k_num_params: gain, mix, three masters, warp, phase,
freq/res/lp × 5) ride identical per-sample linear ramps. The bank keeps one
count, m_ramps_active, and one flag, m_derived_dirty; process()
advances only live ramps and recomputes the derived values (taps, fb,
coefficients, mix gains) per sample while anything moves, once when
everything has settled — the same two-tier idea as svf.h's coefficient
cache, so steady state pays nothing for the smoothness.
The 16-slot preset morph is the same machinery pointed at all 22 targets at
once: store_preset() snapshots the ramp targets (knob positions, not
mid-ramp values); recall_preset(slot, seconds) calls ramp_to() on every
parameter over n = seconds·m_sr samples. Because ramp_to() always
retargets from r.current, a recall issued mid-morph — or a single slider
grabbed mid-recall — is continuous by construction; no special case exists.
The scenarios pin it: a 50 ms recall under a running sine produces no
sample-to-sample jump above the click threshold and lands every parameter
exactly on the preset; a mid-morph set_freq() reaches its own value while
the other 21 keep morphing.
The engineering ledger
k_fb_max = 0.99999, applied after the rt60 solve. The cap makes "res 100 = longest possible resonance" a bounded statement; the rt60 math would happily request fb ≥ 1 for rt60 → ∞ (andres_mastercan push res_eff to 200).- Anti-denormal guard (
|x| < 1e−15 → 0, thetap.comb~idiom) on every recursive state — a comb ringing into silence otherwise decays into denormal territory and multiplies its CPU cost at the quietest moment. - Wet 1/5 normalization, a deliberate deviation.
k_wet_norm = 0.2scales the five-voice sum; the legacy abstraction wired fivetap.comb~objects straight into the output gain — a hot sum, +14 dB at full wet. The mix scenarios pin both endpoints: mix 0 is an exact passthrough, and mix 100 with all resonance off passes at unity to 1e−9. - Equal-power mix:
m_dry_gain = cos θ · g,m_wet_gain = sin θ · g · k_wet_norm, θ = mix·π/200 — matchingtap.crossfade~. - The pickup tap is linear-interpolated (
read_linear), not Hermite — deliberate asymmetry: it is a feed-forward output tap, so its droop is applied once, not compounded per round trip; the argument that disqualified linear ford_readdoesn't apply. - Allocation only in
prepare(): oneceil(sr/k_freq_floor_hz)+8buffer per voice (the 5 Hz floor honors the legacy 200 ms buffer). Every setter is a clamp plus a ramp retarget — safe while audio runs.
Checkpoint
A feedback comb is a string, and the file's one structural opinion is that
the string's decay time — not its feedback coefficient — is the musical
truth: fb = 10^(−3·d/(rt60·fs)), re-solved from the current delay so pitch
sweeps preserve the ring. Hermite taps make the tuning real, the exact
one-pole makes the damping knob honest, and the DC blocker retires the
clipper by making high resonance a solved inequality (fine print stated)
instead of a distortion stage. Warp is pure phase — dispersion paid back to
the fundamental through allpass_phase_delay — and phase is pure geometry:
half a loop is nπ, evens cancel, odds double. The rest is 22 ramps and one
dirty flag, and the tests hold every claim.
Grains that sum to one: grm_pitchaccum.h
The user-facing chapter sold the effect on one image —
+7 becomes +14 becomes +21, a staircase — and one engineering promise: "the
tenth pass is as steady as the first." The image is a topology claim and the
promise is an identity about a pair of window functions, and both are
provable. This appendix proves them with the file's own names, then walks
the pitch follower's failure mode and the ledger. The kernel scenarios in
tap.pitchaccum_tilde_test.cpp drive tap::tools::pitchaccum::accum_bank
directly and pin every measured claim; the sibling tap.shift~ tests pin
the envelope identity to nine decimal places.
Transposition is a moving tap
A delay tap that moves changes pitch. Write the read position of a tap with delay D(n) samples behind a write head at sample n:
p(n) = n − D(n)
The output at sample n reproduces the input's phase at time p(n), so the output advances through the input at the rate:
dp/dn = 1 − dD/dn
A fixed tap (dD/dn = 0) plays at unity. A tap whose delay shrinks by
(ratio − 1) samples per sample plays the buffer at ratio times real speed
— up an octave means eating the delay line at one extra sample per sample.
The code implements exactly this, inverted into phasor form: the tap's delay
is base + window_samples · ph, and the phasor steps
m_phase += −(ratio − 1.0) / window_samples
so dD/dn = window_samples · dph/dn = −(ratio − 1), giving dp/dn = ratio.
(The tt_shift provenance of that line is flagged in the code.) The
transposition scenarios pin the result at the endpoints and the middle:
+12 puts the energy of a 440 Hz sine at 880 Hz, −12 at 220 Hz, and 0 passes
440 untouched — each dominating its reference bin by the test's margins.
One moving tap cannot run forever — the phasor wraps, and at the wrap the
tap teleports across the window: a splice. Hence the classic two-tap engine:
a second tap rides the same phasor at ph_b = ph_a + 0.5 (mod 1), half a
cycle apart, so one tap is always mid-window while the other is wrapping,
and each is faded by envelope() so the splice happens at zero gain.
The envelope pair: an exact partition of unity
Here is envelope(ph, flank), region by region (flank ∈ (0, 0.5] is the
crossfade width as a fraction of the cycle — m_flank maps xfade 1–100%
onto (0.005, 0.5]):
ph ∈ [0, flank]: sin²( π·ph / (2·flank) ) — cos²-shaped rise
ph ∈ (flank, 0.5]: 1 — plateau
ph ∈ (0.5, 0.5+flank]: cos²( π·(ph−0.5) / (2·flank) )— fall
ph ∈ (0.5+flank, 1): 0
The claim — stated as a comment in the file, and load-bearing — is that the taps at ph and ph + 0.5 sum to 1 exactly, at every phase and every flank width. Proof by the same four regions, writing e(·) for the envelope and using ph_b = ph_a + 0.5 mod 1:
ph_a ∈ [0, flank]: ph_b ∈ [0.5, 0.5+flank]
e_a + e_b = sin²(π·ph_a/2·flank) + cos²(π·ph_a/2·flank) = 1
ph_a ∈ (flank, 0.5]: ph_b ∈ (0.5+flank, 1]
e_a + e_b = 1 + 0 = 1
ph_a ∈ (0.5, 0.5+flank]: ph_b wraps to ph_a − 0.5 ∈ (0, flank]
e_a + e_b = cos²(π·(ph_a−0.5)/2·flank) + sin²(π·(ph_a−0.5)/2·flank) = 1
ph_a ∈ (0.5+flank, 1): ph_b = ph_a − 0.5 ∈ (flank, 0.5]
e_a + e_b = 0 + 1 = 1
Every crossfade region pairs a sin² with the cos² of the same argument; everywhere else a plateau pairs with a zero. The construction also joins each flank to its plateau with zero slope (the derivative of sin² vanishes at both ends), so there is no corner to click, and at flank = 0.5 the plateau vanishes and the pair degenerates to a complementary Hann pair.
Why demand exact? In a one-shot shifter, an envelope pair that sums to
1 ± δ is a gain ripple of δ at the grain rate — a subtle tremolo, mildly
regrettable. But this object's entire identity is that the transposer sits
inside a feedback loop: the ripple multiplies onto the signal on every
pass, and after k trips the peaks have compounded to (1+δ)ᵏ while the
troughs have decayed — a pumping that grows with exactly the feedback
settings the effect is played at. That is why the original engine's window
was replaced: tt_shift used a fixed 256-point padded-Welch table whose
pair did not sum flat (the deviation is documented in both this file's
header and tap.shift~'s). The identity is pinned numerically in the
tap.shift~ wrapper tests — same envelope construction at flank = 0.5 —
where DC at 0.5 pushed through moving taps at ratio 1.3 comes out equal
to the input with max error below 1e−9: no grain-rate ripple, to double
precision.
The accumulation topology
transposer::process() wires the loop in this order:
in ──►(+)──► delay buffer ──► two moving enveloped taps ──► y (out)
▲ │
└── × fb ◄── DC blocker ◄── m_fb_state (previous y) ◄─┘
The fed-back sample re-enters upstream of the taps: it is written into the
buffer and then read back by the moving, windowed taps — which is to say it
is delayed by delay_samples and transposed by ratio again. Every
trip multiplies the frequency by another factor of ratio: +7 st becomes
+14 becomes +21. Contrast the ordinary "feedback around a shifter" patch,
where the feedback taps the delay output and re-enters the delay: each
echo is re-delayed but shifted only once, and the staircase never climbs.
The topology is the effect, and the kernel test pins its two-pass
signature: a 440 Hz burst at +7 st, 300 ms delay, 70% feedback shows
Goertzel energy at 659.26 Hz (one pass) in the 0.32–0.55 s window and at
987.77 Hz (two passes — the accumulation itself) in the 0.65–1.0 s window,
each dominating the off-frequency reference bin.
Boundedness: the loop gain really is fb
The constant-sum envelope has a second payoff. Since e_a, e_b ≥ 0 and
e_a + e_b = 1, the two-tap read is a convex combination of past loop
samples — with linear taps its magnitude could never exceed the buffer's
peak, and the Hermite taps can overshoot that bound only by the
interpolator's small, bounded ripple. So the per-pass gain around the loop
is genuinely fb — capped at k_fb_max = 0.99 — and not fb × (envelope
ripple peak), which is the number that would have mattered with an uneven
pair. The DC blocker in each loop kills the one component granular splicing
can otherwise rectify into a ramp; its own tiny high-frequency shelf
(2/(1+R) ≈ 1.0005) is absorbed a hundred times over by the 1% headroom in
the cap — this kernel does not need the fine print that grm_comb.h's
0.99999 cap does. The test drives the worst case: both voices at 99%
feedback, opposing transpositions, five seconds — output finite, peak < 50.
Modulation: one LFO, two seeded dice
Per sample, each voice's transposition is assembled as
trans_eff = m_ramp[p_trans1 + v].current + lfo + rnd
ratio = 2^(trans_eff / 12)
The LFO is one global phasor (m_lfo_phase); voice 2 reads it at
+ modphase/360, so the two shadows breathe against each other at a
settable phase — 90° by default. The random component is per-voice:
tick_random() holds a target from a linear-congruential generator
(seeded 1111 and 2222 in the constructor) and cosine-interpolates between
held values at randrate, re-arming its phase when disabled so re-enabling
starts a fresh segment rather than finishing a stale one. Deterministic
seeding is a testability decision that is also a musical one — the same
patch renders the same audio — and the test takes it literally: two
identically configured banks render outputs with maxdiff == 0.0, bit
identical. The modulation scenario pins the spectral effect from the other
side: 1 st of 5 Hz LFO drains more than half the carrier bin's energy into
sidebands.
The pitch follower, and the subharmonic trap
follow adapts the grain window to the input. The follower is deliberately
cheap: input decimated by k_flw_decim = 8 (6 kHz at 48 kHz) into a
1024-float ring, and every k_flw_interval = 512 input samples a normalized
autocorrelation over a k_flw_win = 512 window:
corr[lag] = Σ x[n]·x[n+lag] / √( Σ x[n]² · Σ x[n+lag]² )
searched over lags for 50–800 Hz. The naive readout — take the global
maximum — has a classic failure mode that the code's comment names: a
periodic signal correlates at every multiple of its period, so corr[2T]
and corr[3T] sit at essentially the same height as corr[T], and windowing
noise routinely pushes a multiple to the numerical top. The global argmax
then reports a subharmonic — an octave or twelfth low — and the window
snaps to twice the true period. The fix is two lines: find the global max
best_r, then take the smallest lag whose correlation clears
accept = max(k_flw_confidence, 0.85·best_r) — the earliest lag within 15%
of the peak. Below k_flw_confidence = 0.6 nothing is accepted and
period_s() reports 0: unpitched input is ignored rather than chased.
The window goal is k_flw_periods = 2 detected periods (clamped to the
5–200 ms window range), approached through a one-pole slew
(m_window_eff_ms += 0.0005·(goal − eff), ≈ 40 ms time constant at 48 kHz)
that relaxes back to the manual window value when follow is off or the
gate says unpitched. Why two periods: at each grain wrap the tap jumps by
exactly the window, so a window of k whole periods makes the splice
displacement an integer number of cycles — the spliced waveform stays
phase-coherent and the grain-rate artifacts land in tune with the source
instead of at an arbitrary rate (the envelope cycle rate is
|ratio − 1|·fs/window_samples, which at W = 2T scales with the source pitch,
≈ (ratio−1)·f₀/2). And k = 2 rather than 1 keeps the window above the 5 ms
floor across most of the follower's 50–800 Hz range. The test pins both the
adaptation and the trap: a 220 Hz tone converges the effective window into
6–13 ms, bracketing two periods (9.1 ms) — a subharmonic lock would demand
18.2 ms, outside the pin — and white noise leaves the window at the manual
setting (70–100 ms around the 87 ms default).
The engineering ledger
- 17 ramped parameters, same morph engine as
grm_comb.h. One ramp array,m_ramps_active,m_derived_dirtyrecomputing derived values per sample while moving and once at settle;store_presetsnapshots targets,recall_presetretargets every ramp from its current value, so mid-morph overrides are continuous with no special case. Pinned: an 80 ms recall under audio shows no sample-to-sample jump above 0.3 and lands every parameter to 1e−9; mix 0 is an exact passthrough (< 1e−9) even with 90% feedback churning inside the muted wet path. - The modulated ratio path lives in
process(), notupdate_derived()— the code comments the split: LFO and random are inherently per-sample, so caching them would save nothing; the cacheable tier is delays, fb, voice gains, flank, and the equal-power mix. - Hermite taps with
k_base_delay = 3headroom — same 4-point interpolator and same ≥ 2-sample geometry constraint asgrm_comb.h, here so that a voice delay of 0 ms is still legal under moving taps. - GRM's stereo-width fader is dropped, on purpose. The kernel is mono
and the wrapper is single-channel by house rule (
mc.wraps it); the omission is declared in the wrapper header and the maxref. Width would have been the only parameter that could not live in a mono kernel. - The follower is a mode, not a fader —
set_follow(bool)sits outside the morphable parameter set (the file says so), because interpolating a boolean analysis mode over a morph is meaningless. - Allocation discipline. One buffer per voice sized for the worst case
(
(k_max_delay_ms + k_max_window_ms)at prepare-time sr, +16), a fixed 1024-float follower ring, and a stack-local correlation array; afterprepare()the audio path allocates nothing, and the analysis cost — a few hundred multiplies per input sample, amortized — is paid only whenfollowis enabled.
Checkpoint
A tap whose delay changes at (1 − ratio) samples per sample is a transposer — dp/dn = ratio, by one derivative. Two taps on the same phasor half a cycle apart cover each other's splices, and the cos²/sin² envelope pair sums to one exactly, region by region, at any flank width — which in a feedback loop is the difference between loop gain fb and loop gain fb-times-ripple compounding every pass. The feedback re-enters upstream of the taps, so every echo is transposed again: the staircase is a topology, and the test hears both steps. The follower reads the earliest strong autocorrelation lag, not the tallest, because the tallest is routinely a subharmonic; two detected periods keep the splices phase-coherent. Everything else — ramps, morph, seeds, caps — is the same discipline as the comb bank, and equally pinned.
Two banks and a multiplier: vocoder.h
The user-facing chapter made three flat promises about
tap.vocoder~: a silent carrier is silence, gain is exactly linear, and a
silent modulator decays away at the follower rate. It could afford to,
because none of those is a tuning outcome — each one is a structural fact
about a very small graph. This appendix draws the graph, proves the facts,
and then walks the three numerical choices (band placement, filter type,
follower coefficient) that make the graph sound like a vocoder.
One honesty note up front. The original tap.vocoder~ source did not
survive the revival; vocoder.h is reconstructed from the reference
documentation — "a basic 24-band vocoder" with q and response_interval
attributes. The topology below is the classic channel vocoder that
documentation describes, and the tests pin its structural behavior; there is
no lost binary to bit-compare against, and this chapter never pretends
otherwise.
The graph: a bilinear form in 24 subbands
A channel vocoder is subband multiplication. Split both signals with the same filter bank, measure the modulator's level per band, scale each carrier band by that level, sum:
band i: m_i = B_i(modulator) the modulator through bandpass i
env_i ← follower(|m_i|) its envelope
c_i = B_i(carrier) the carrier through the identical bandpass
output: y = gain · Σᵢ c_i · env_i
That is bank::process() verbatim — the loop body computes m, rect,
m_env[i], c, and accumulates c * m_env[i], and the return line applies
m_gain once to the sum. Three contracts follow from the shape alone, and
tests/vocoder_test.cpp pins each one:
- Silent carrier ⇒ exactly silence. Every summand carries a factor
c_i. A biquad is linear with zero state at rest, so a zero carrier givesc_i ≡ 0for all i, and the sum is identically zero no matter what the modulator (and hence the envelopes) does. The test drives a 220 Hz modulator against a zero carrier for 8000 samples and requires peak < 10⁻¹², but the true bound is exact:0.0 * m_env[i]is 0.0. - Gain is exactly linear. The multiply
c_i · env_iis the only nonlinearity in the graph, and it is bilinear — linear in the carrier with the envelopes held fixed, linear in the envelopes with the carrier held fixed.m_gainsits outside all of it, a scalar on the finished sum, and nothing upstream reads it. Two banks fed identical inputs with gains 1 and 2 must differ by exactly a factor of 2, float for float; the test requires|yb − 2·ya| < 10⁻¹²across 8000 samples. - Silent modulator ⇒ output decays at the follower rate. With the
modulator silenced,
rect = 0and each envelope obeysenv ← m_env_coef · env— a geometric decay with the follower's time constant. The output is bounded byΣ|c_i|·env_i, so it decays with the envelopes even while the carrier keeps playing. The test warms the bank up, silences the modulator for one second at 48 kHz (≈ 50 time constants at the 20 ms default — a decay of e⁻⁵⁰), and requires the late output under 10⁻⁴ of the warmed level.
The fourth pinned property, determinism (two identical runs compare equal
with ==), is the repo-wide claim that the kernel is pure state-machine
arithmetic: no randomness, no time, no allocation in the audio path.
The bilinear form in 24 subbands — the graph shape the proofs read off.
Where the bands sit
Twenty-four bands span 50 Hz to 12 kHz, log-spaced. band_frequency(i)
computes:
f_i = k_fmin · (k_fmax / k_fmin)^(i / (k_bands − 1)) i = 0 … 23
= 50 · 240^(i/23)
so adjacent centres sit at a constant ratio of 240^(1/23) ≈ 1.269 — about
0.344 octave, a hair over four semitones, per band. Log spacing is the only
defensible choice for this machine, twice over: the ear judges musical width
by ratio, not by hertz, so equal-ratio bands devote equal perceptual width
to each channel; and speech puts its identity (formants, the envelope the
vocoder exists to capture) in the low kilohertz while its detail
(fricatives) rides above — a linear spacing would waste twenty bands above
6 kHz and cram every vowel into two. The span itself brackets speech: 50 Hz
is below any voice fundamental, 12 kHz is above any formant that matters,
and recalc_filters() clamps each centre at 0.45 · m_sr so the top bands
stay well below Nyquist at low sample rates rather than folding.
The filter: constant peak, unconditional stability
Each band is an RBJ Audio-EQ-Cookbook bandpass, the constant 0 dB-peak
variant, computed in recalc_filters():
w0 = 2π · fc / sr alpha = sin(w0) / (2·q) a0 = 1 + alpha
b0 = alpha / a0 a1 = (−2·cos w0) / a0
b1 = 0
b2 = −alpha / a0 a2 = (1 − alpha) / a0
"Constant 0 dB peak" is a normalization claim: the gain at the centre frequency is exactly 1, for any Q. It is worth proving, because the whole level architecture rests on it. Evaluate the transfer function at z = e^(jw0):
H(z) = alpha·(1 − z⁻²) / [(1 + alpha) − 2·cos w0 · z⁻¹ + (1 − alpha)·z⁻²]
denominator at z = e^(jw0):
[1 − 2·cos w0 · e^(−jw0) + e^(−2jw0)] + alpha·(1 − e^(−2jw0))
= e^(−jw0)·(e^(jw0) − 2·cos w0 + e^(−jw0)) + alpha·(1 − e^(−2jw0))
= e^(−jw0)·(2·cos w0 − 2·cos w0) + alpha·(1 − e^(−2jw0))
= alpha·(1 − e^(−2jw0)) = the numerator exactly
so H(e^(jw0)) = 1 identically. Why it matters here: env_i is supposed to
measure the signal's level in band i, and each carrier band is supposed to
be scaled by that measurement and nothing else. With the constant-peak
variant, changing q changes bandwidth only — the on-centre gain of all 48
filters stays pinned at unity, so the q knob narrows or overlaps the bands
without re-balancing the reconstructed spectrum or re-calibrating the
envelope levels. The cookbook's other bandpass (constant skirt gain) has
peak gain Q; with the default q = 20 that would be +26 dB per band, scaling
with the knob — every q move would also be a 24-band gain move.
Stability is likewise unconditional. A biquad is stable iff its
coefficients sit in the stability triangle, |a2| < 1 and |a1| < 1 + a2.
Here a2 = (1 − alpha)/(1 + alpha), which lies in (−1, 1) whenever
alpha > 0 — and alpha = sin(w0)/(2q) is positive for any q > 0 and any
0 < fc < Nyquist; the second condition, 2|cos w0|/(1 + alpha) <
1 + (1 − alpha)/(1 + alpha) = 2/(1 + alpha), reduces to |cos w0| < 1, true
on the same range. The code enforces the preconditions rather than
assuming them: q is floored at 0.001 and fc clamped to 0.45·sr, so no
attribute value and no sample rate can produce an unstable band. The
sections run as Direct Form I (biquad::process keeps x1, x2, y1, y2) —
at these moderate Qs and double precision, the plainest form is the
honest one.
The follower: one coefficient, symmetric by construction
Each band's envelope is a one-pole lowpass over the full-wave rectified band signal:
rect = |m_i|
env_i ← m_env_coef · env_i + (1 − m_env_coef) · rect
with the coefficient computed in recalc_envelope() from the
response_interval attribute:
tau = response_ms / 1000 (ms → seconds)
m_env_coef = exp(−1 / (tau · sr))
That is the exact one-sample step of a continuous first-order lag with time
constant τ: the discrete pole e^(−T/τ) with T = 1/sr. So the documented
"analysis period" is a time constant, precisely — after response_interval
milliseconds of silence an envelope has decayed to 1/e of its value, and
after a step up it has covered 1 − 1/e of the distance. Note what the code
does not have: separate attack and release. One coefficient serves both
directions, which is what the legacy surface documents (a single
response_interval) and is why the user chapter calls the knob "the
vocoder's attack and release." The 10⁻⁴ floor on response_ms keeps the
exponent finite; at the 20 ms default and 48 kHz, m_env_coef ≈ 0.99896.
Why time-domain, when the siblings went spectral
tap.nr~ and tap.spectra~ (next chapter) are STFT machines. The vocoder
deliberately is not, for three compounding reasons:
- Zero algorithmic latency. The spectral scaffold costs exactly one FFT frame of delay by construction; this graph's output at sample t depends only on inputs up to t. A vocoder is played live against its carrier — latency is a musical defect here in a way it is not for noise reduction.
- It is cheap. 48 biquads (5 multiplies + 4 adds each in DF I) plus 24 follower updates and 24 multiply-accumulates — on the order of three hundred flops per sample, no transform, no windowing, no frame buffers.
- It is faithful. The original
tap.vocoder~was a real time-domain external; the pfft~-hosted abstraction that wrapped it in some patches only added smoothing and gain around it. Rebuilding it as an FFT effect would have been reconstructing a different object. Sovocoder.hfollows thesvf.h/ladder.hidiom —prepare(samplerate)then per-sampleprocess()— not theconfigure(fftsize)scaffold of the spectral set.
The engineering ledger
prepare()recomputes everything. It callsrecalc_filters()(24 coefficient sets, each written into bothm_mod[i]andm_car[i]— the banks are identical by construction, one computation assigned twice) andrecalc_envelope().set_qre-runs only the filters,set_response_msonly the envelope coefficient,set_gainis a bare store — each setter pays for exactly what it moves, the small-scale version ofsvf.h's two-tier update.- Setters are allocation-free and audio-safe. All state is in fixed
std::arrays sized byk_bands; there is no allocation anywhere in the class, so the Min wrapper can forward attribute changes from the message thread while the perform loop runs. - The legacy surface is honored, with one documented fix.
qandresponse_intervalkeep their documented names, meanings, and defaults (20 and 20 ms). The original registered both attributes assymbol; the wrapper (tap.vocoder_tilde.cpp) registers them asnumber, which is what they actually are — a Q value and a millisecond time — and says so in its header.gainis a small, admitted addition for level staging, since a band-multiplied signal lands quieter than either input. - Both banks clear together.
clear()zeroes all 48 biquad states and the envelope array — the whole graph's memory, nothing else, so aclearmessage can never leave a stale envelope gating a fresh carrier. - What is deliberately absent: per-band gain trims, separate attack/release, a noise-driven "unvoiced" band — all classic vocoder extensions, all outside the documented surface being reconstructed. The reference page promised a basic 24-band vocoder; the file implements exactly that and stops.
Checkpoint
The vocoder is a bilinear form: two identical 24-band banks and one multiply per band. Everything the tests pin — silence in, silence out; exact gain linearity; follower-rate release — is a consequence of that shape, not of tuning. The numerics are three choices: log spacing (equal ratio per band, matched to hearing and to speech), the constant-peak RBJ bandpass (band level measures the signal, not the Q, and stability is a theorem with the clamps in place), and the exact one-pole coefficient e^(−1/(τ·sr)) (the documented period is an honest time constant, symmetric in both directions). Time-domain because latency, cost, and history all point the same way.
One STFT, three effects: fft.h, stft.h, nr.h, spectra.h
The user-facing chapters for tap.nr~ and
tap.spectra~ both lean on the same claim: the machinery
is transparent — set the effect to do nothing and the output is the input,
exactly, one FFT frame late. All the trust in these objects lives in that
claim, and it is not free: it has to be engineered into the window, the
overlap, and one normalization constant. This appendix builds the stack
bottom-up — the FFT, the scaffold, then the two small effects on top — and
proves the transparency claim rather than asserting it.
fft.h: the transform, owned outright
The kernel repo's law is zero dependencies — plain C++17, standard library
only. So the FFT is in-house: an in-place iterative radix-2 Cooley–Tukey
in fft::transform(re, im, inverse), about forty lines. It is also one
copy by design: the identical routine previously lived, byte for byte,
inside conv_engine (tap.convolve~), tap.nr~, and tap.spectra~, and
was consolidated at the kernel split so it is maintained and tested once —
tests/fft_test.cpp is that single test point. The two halves:
- Bit-reversal permutation. An iterative FFT consumes its input in
bit-reversed index order; the first loop swaps each element
iwith its bit-reversed partnerj, maintainingjincrementally (the carry-ripple idiom) rather than reversing bits per index; thei < jguard swaps each pair once. - Butterfly stages. For each length
len= 2, 4, … N, combine pairs of half-blocks with twiddle factors e^(∓2πik/len). The twiddle is advanced by a complex-multiply recurrence (cwr,cwirotated by (wr,wi)) — onecos/sinper stage instead of per butterfly. A recurrence accumulates rounding, but in double precision over these sizes it is far inside the pinned tolerance.
The scaling convention is asymmetric and load-bearing: forward is
unscaled, inverse divides by N, so forward-then-inverse is the identity.
Every claim is pinned in fft_test.cpp: the forward transform matches a
naive O(N²) DFT to 10⁻⁹ for N ∈ {2, 4, 8, 16, 64, 256}; the round trip
reconstructs random complex input to 10⁻⁹ at N = 128; a unit impulse
transforms to an exactly flat unit spectrum; and a real cosine at bin 3 of
32 lands N/2 on bins 3 and 29 — fixing the sign convention (forward kernel
e^(−i…)) and the conjugate-bin layout the effects below depend on.
stft.h: the scaffold, and the COLA proof
One pump, two effects: nr and spectra are this pipeline with different middles.
stft is the overlap-add machinery shared verbatim by both effects: Hann
window, fixed 4× overlap (m_hop = m_fftsize / m_overlap), a circular
input buffer, a circular output accumulator, and a per-sample pump.
process() takes the effect as a callable — op(re, im, N) mutates the
N-point spectrum in place between the forward and inverse transforms; the
only difference between tap.nr~ and tap.spectra~ is that lambda. The
window, built in configure():
m_window[k] = 0.5 − 0.5·cos(2π·k / m_fftsize) k = 0 … N−1
— the periodic Hann (denominator N, not N−1), which is what makes the
overlap sums below exactly constant rather than rippling. The window is
applied twice per frame: once at analysis (m_re[k] = inbuf·window[k])
and once at synthesis (outbuf += m_re[k]·m_window[k]·m_norm). With an
identity op, the inverse transform returns the windowed frame exactly (the
FFT round trip is the identity), so each input sample x is delivered to the
output through every frame that covers it, weighted by w² each time:
y[t] ∝ x[t−N] · Σₘ w²(n − mH) H = N/4, four frames cover each n
Perfect reconstruction therefore requires the shifted window-squared sum to be constant — the COLA (constant overlap-add) condition for double-windowing. For the periodic Hann at 4× overlap it is, and the constant has a closed form. Expand w²:
w²(θ) = (0.5 − 0.5·cos θ)² = 0.375 − 0.5·cos θ + 0.125·cos 2θ θ = 2πn/N
A hop of N/4 advances θ by π/2. Over four hops, the cos θ terms are four quarter-turns of a phasor — they sum to zero; the cos 2θ terms advance by π per hop and cancel in adjacent pairs. What survives is the constant:
Σₘ w²(n − mH) = 4 × 0.375 = 3/2 for every n
The code does not hard-code 3/2. configure() overlap-adds overlap
copies of m_window[k]² around a circular buffer and reads the value at
cola[m_fftsize/2], setting m_norm = 1/c — for Hann at 4×,
m_norm = 2/3 (verified numerically: the computed sum is 1.5 to within
10⁻¹⁵ at every index, so reading the midpoint is safe). Computing it keeps
the scaffold correct for any window/overlap it might grow.
Latency is exactly N, and here is the accounting. The pump writes
in[i] into m_inbuf[m_pos], reads the output from m_outbuf[m_pos],
zeroes that slot, advances, and fires a frame every m_hop samples. When a
frame fires, its index k holds input sample x[t₀ − (N−1) + k] (t₀ the
newest sample, at k = N−1), and synthesis writes index k into
m_outbuf[(m_pos + k) % N], which the pump reads k+1 samples later. Output
time minus input time:
(t₀ + 1 + k) − (t₀ − N + 1 + k) = N for every k, every frame
A frame cannot be transformed until it has filled — that is the whole cost,
and why latency() simply returns m_fftsize. Both test suites pin the
full contract at once: with a do-nothing effect (nr at threshold 0,
spectra at remap 1), out[t] == in[t − N] to within 10⁻⁹ on broadband
noise for all t ≥ 2N (the run-in covers frames that still window in
zeros). This is the "transparent machinery" sentence in both user
chapters, with its provenance attached: FFT round trip (pinned) × COLA
constant (derived) × exact-N pipeline (derived).
nr.h: the gate, precisely
The spectral op is gate(), and its knee is short enough to quote in full
as math. Per bin k:
mag = √(re[k]² + im[k]²) · (2/N)
gain = 1 if thr ≤ 0 or mag ≥ thr
gain = (mag / thr)^slope if mag < thr (slope ≤ 0 → 1)
re[k] *= gain; im[k] *= gain
The 2/N puts mag on a sinusoid-amplitude scale (a real tone of amplitude
A puts A·N/2 in each of its two conjugate bins; ×2/N recovers A). One
honest calibration note: the frame is Hann-windowed before the FFT, and
the Hann's coherent gain is 1/2 — a bin-centred sine of amplitude A
actually measures mag = A/2 (verified numerically: A = 0.8 reads 0.400),
with leakage in the adjacent bins. threshold is a linear amplitude on the
windowed scale; a full-scale sine sits near 0.5, not 1.0.
The knee is a downward expander per bin. Take logs of the gain law below threshold:
L_out − L_thr = (1 + slope) · (L_in − L_thr) in dB
Every dB below the threshold becomes (1 + slope) dB below it: slope 0 is
unity (bypass by another name — the code special-cases it), the default
slope 2 is a 1:3 expander, and slope → ∞ approaches a hard gate. Both
re and im are scaled by the same real gain, so phase is untouched — the
gate reshapes magnitude only.
Two structural notes. First, the loop runs over all N bins, mirror
half included, with no symmetry bookkeeping — and needs none: the input
frame is real, so its spectrum is Hermitian, magnitudes are symmetric
(mag[N−k] = mag[k]), conjugate pairs get the same real gain, and
Hermitian symmetry survives the op — the inverse stays real for free.
(Hold that thought; spectra is not so lucky.)
Second, the per-frame independence of the gain decision is exactly where
musical noise comes from: a bin whose magnitude hovers near thr flips
between pass and heavy attenuation frame by frame, and each isolated pass
is one Hann-windowed near-sinusoid burst, milliseconds long — a chirp.
Scattered over time and frequency, chirps sound like water. That is not a
bug in the code; it is the knee's steepness meeting the frame rate, which
is why the user chapter's cure is a gentler slope, not a different
implementation.
tests/nr_test.cpp pins the three defining behaviors: gate open
(threshold 0) reconstructs noise to 10⁻⁹ delayed one frame; a quiet
bin-centred tone (amplitude 0.05 against threshold 0.5, slope 4) leaves
a steady-state tail under 5 % of the input RMS; a loud tone (0.8 against
0.01) passes with RMS within 2 % of the input.
spectra.h: the remap, and why the mirror is not optional
The op builds a new spectrum over the lower half:
src = lround(k · m_remap) k = 0 … N/2
m_ore[k] = re[src], m_oim[k] = im[src] if 0 ≤ src ≤ N/2, else 0
then forces DC and Nyquist real (m_oim[0] = m_oim[half] = 0), mirrors —
m_ore[N−k] = m_ore[k], m_oim[N−k] = −m_oim[k] for 0 < k < half — and
copies the scratch back over re/im.
The mirror is provable necessity, not tidiness. A real signal's DFT
satisfies X[N−k] = conj(X[k]), and only Hermitian spectra invert to real
signals. The remap fills the lower half by an arbitrary rule and touches
nothing above Nyquist — the upper half still holds the input's bins, so
the assembled spectrum is not Hermitian for any remap ≠ 1, and its inverse
transform is genuinely complex. And the scaffold's synthesis reads m_re
only — the imaginary part of the inverse is discarded. Keeping
Re(IDFT(Y)) is algebraically inverse-transforming the Hermitian average
½(Y[k] + conj(Y[N−k])): without the mirror, the delivered effect would be
an uncontrolled blend of the remapped lower half and the untouched upper
half, half the intended signal shunted silently into the discarded
imaginary part. The mirror makes the spectrum Hermitian by construction,
so the inverse is exactly real and "keep the real part" loses nothing. DC
and Nyquist are their own mirror images (k = N−k), so conjugate symmetry
forces them real — hence the two explicit zeroes.
(The remap cannot run in place — output bin k may read a bin already
overwritten — hence the m_ore/m_oim scratch, allocated in
configure().) And the energy honesty: a remap is not a permutation, so Parseval is
deliberately broken. For remap < 1, lround(k · remap) is non-strictly
increasing — several output bins read the same input bin, duplicating
its energy. For remap > 1 it strides — input bins are skipped, and every
output bin above N/(2·remap) reads beyond Nyquist and is zeroed,
discarding the input's top octaves. Neither direction conserves energy,
and neither is meant to: the reference page has called this an
"ultra-non-linear effect" since 2002; the kernel implements the rule, not
a transform.
tests/spectra_test.cpp pins the two anchors: remap 1 reconstructs noise
to 10⁻⁹ delayed one frame (the identity copies the lower half of an
already-Hermitian spectrum, and the mirror rebuilds the upper half it
started with); remap 2 moves a tone at input bin 16 to output bin 8 —
lround(8 · 2) = 16 — verified by FFT-ing a frame-aligned slice of the
steady-state output and requiring the peak at bin 8.
The engineering ledger
- The effect is a template parameter, not a base class.
stft::processtakesSpectralOp&&and calls it once per hop; each effect passes a capturing lambda. No virtual dispatch in the audio path, full inlining. - Why the vocoder is not the third client.
tap.vocoder~is time-domain on purpose — zero latency, no frame,prepare(sr)instead ofconfigure(fftsize)— see its own chapter. The spectral set accepts latency as a cost model; the vocoder's whole point is not paying it. - Allocation at
configure()only. Window, in/out rings, FFT scratch, and (forspectra) the remap scratch are all sized there;process()is allocation-free.reset()flushes the running buffers without reallocating or touching the window — commented in the code as safe from a message handler, which is exactly how the wrappers use it. fftsizeis the one shared dial, and the scaffold makes its price explicit: resolution (bin spacing sr/N), smearing (a per-bin decision spreads over a whole frame), and latency (latency()returns N so the wrapper can report a true number to the host).- One FFT, tested once. The round-trip and DFT-reference pins in
fft_test.cppare what let this chapter treat "forward then inverse is the identity" as a premise everywhere above.
Checkpoint
The stack is three honest layers. The FFT is forty owned lines with an
asymmetric scaling convention, pinned against a naive DFT. The scaffold
windows twice, so reconstruction needs the shifted w² sum to be constant —
for periodic Hann at 4× overlap it is exactly 3/2, measured rather than
assumed, and the pipeline delays every sample by exactly N. On top, each
effect is one spectral op: nr a per-bin 1:(1+slope) downward expander
whose real gain preserves Hermitian symmetry for free; spectra a
bin-index rule violent enough that reality — a real output — must be
restored by explicit mirror. Transparency at the neutral setting is the
theorem the whole stack exists to satisfy; both suites pin it at 10⁻⁹.
Seventeen, not four: diode_ladder.h
The transistor-ladder appendix derived why the Moog loop oscillates at
k = 4. The 303's filter looks like the same idea — four capacitors, one
feedback path — and behaves like a different species, because the diode
ladder deletes the one luxury the Moog circuit has: buffering. Every diode
pair both charges the next capacitor and loads the previous one. This
appendix derives what that coupling does to the poles, why the oscillation
threshold lands at exactly 17, why the shipping filter still refuses to
self-oscillate at stock settings, and how the coupled nonlinear system is
solved in closed form every sample.
Everything here is verified against Tim Stinchcombe's published TB-303
circuit analysis and the executed
tb303.ipynb
notebook, which matches the kernel's linearized response to his transfer
function to 0.028 dB.
The chain: a diffusion line, not a cascade
With the diode conduction curve linearized (identity for now), the four node voltages obey a coupled chain:
v1' = ω·(S(u − v1) − S(v1 − v2))
v2' = ω·(S(v1 − v2) − S(v2 − v3))
v3' = ω·(S(v2 − v3) − S(v3 − v4))
v4' = 2ω·S(v3 − v4) [the top capacitor is halved on the schematic]
Each middle equation has two terms — charge in from the left, charge stolen by the right. That is the loading, and it is the whole story: this is a discrete diffusion line, not four independent one-poles.
Every edge is a tanh; every middle node leaks both ways. The bidirectional arrows are what a buffered cascade doesn't have — and why the poles spread. Its normalized transfer function works out to exactly Stinchcombe's measured TB-303 response,
H(s) = 1 / (s⁴ + 6.727·s³ + 14.142·s² + 9.514·s + 1)
and the coefficients are not arbitrary — they are 4·2^(3/4), 10·√2,
8·2^(1/4): the equal-component chain with the top cap halved, which is
also why Stinchcombe finds that changing that one cap shifts cutoff by
2^0.25. The poles are all real and spread ~25:1 (−0.13, −1.04, −2.33,
−3.24 normalized). Compare the Moog ladder: four coincident poles.
Consequences you can hear:
- Asymptotically the slope is 24 dB/oct, but only ~14 dB falls in the first octave above cutoff — the honest version of the panel's "18 dB" claim.
- At resonance 0 the −3 dB point sits ~3.2 octaves below the resonance
frequency. The kernel's
frequencyparameter names the resonance peak, and the wide skirt below it is the real filter, not a tuning bug.
Seventeen: the closed loop's threshold
Feedback enters as u = drive·x − k·hp(v4). Ignore the high-pass for a
moment and run Routh–Hurwitz on the closed loop: the stability boundary
lands at exactly k = 17, with the marginal oscillation at √2× the
stage rate. (Open303 normalizes its feedback by the same 1/17.) So
resonance maps k = 17·resonance, putting 1.0 at the ideal chain's
threshold — and the prewarp is chosen so that the √2 factor lands the
oscillation on the labeled frequency:
g = tan(π·fc / fs_os) / √2 per stage (2g on the top stage)
Why a stock 303 never quite sings
Now put the high-pass back. The hardware's resonance feedback runs through a
~150 Hz one-pole high-pass (Open303's calibrated value, the fbhp default),
and its phase lead pushes the would-be oscillation frequency up — to where
the ladder attenuates more. Measured on the shipping kernel: the closed loop
needs k ≈ 17.5 even at 8 kHz, ≈ 19 at 2 kHz, ≈ 25 at 500 Hz. The knob
stops at 17. The emergent result — not programmed, derived — is the famous
trait: a stock TB-303 never quite self-oscillates, and neither does this
filter until you take the documented bend (resonance runs to 1.5, i.e.
k = 25.5; past ~1.1 it sings at high cutoffs, slightly sharp of fc for the
same phase-lead reason).
The high-pass buys two more behaviors for free:
- Resonance thins as cutoff falls — low notes squelch instead of ringing, which is why a 303 keeps its bass at high resonance.
- Closed-loop DC gain is exactly 1 regardless of resonance. The
transistor ladder needed a
compparameter to buy its passband back; the 303's own circuit is the compensation, so this kernel has none.
Set fbhp 0 and the ideal analysis becomes exact: threshold at 1.0,
oscillation at fc, drifting flat by ~0.7× the resonance excess past
threshold — an amplitude effect (the growing swing saturates the edges
unevenly), identical at every cutoff, and pinned by test.
The nonlinearity lives on the edges
In the circuit the coupling elements saturate — there are no buffer amps
between stages to saturate instead. So the kernel puts tanh on every
S(·) above: four saturators, one per diode-pair edge, slope 1 at the
origin so small signals see exactly the linearized Stinchcombe response.
There is no asym parameter here, deliberately: the diode pairs are
complementary, so the transistor ladder's operating-point-mismatch story
does not apply.
Solving the coupled system in closed form
The ZDF discretization (trapezoidal, as everywhere in the house) turns each
sample into a system: five unknowns (v1..v4 and the loop input u)
that all depend on each other through the couplings and the feedback. The
kernel linearizes each edge with a secant gain γ = tanh(e)/e at an
operating point, and then — this is the part worth reading in the code —
eliminates the linear system bottom-up to closed form: v4 in terms of
v3, then v3 = p30 + p32·v2, v2 = q20 + q21·v1, back-substituted until
one division yields v1 and everything else follows. No matrix, no pivots,
and unconditionally stable: every divisor is ≥ 1 for g > 0, γ ∈ (0, 1].
The feedback high-pass's state enters the same solve (its instantaneous
gain 1 − G multiplies k), so the loop is closed exactly, high-pass
included.
The two solvers differ only in how the secant gains chase the operating point:
solver_fast(default): solve at the previous sample's gains, refresh the gains at that solution, solve once more, commit.solver_exact: repeat the refresh-and-solve until the node voltages move by less than 1e-12 (capped at 32 iterations).
Measured across a settings matrix out to resonance 1.4 and +24 dB drive —
beyond hardware reach — the worst-case difference is −44.9 dBr, at
1.6–3.3× the CPU. The fast path's one correction is almost always enough
because tanh is smooth and the operating point moves slowly at audio rate;
the exact path exists so that claim never has to be taken on faith.
After the solve, the states advance trapezoidally using the true diode
currents (tanh at the solved voltages, not the secant approximations) —
the same "linearize to solve, commit with the real nonlinearity" pattern as
ladder.h's one-pass commit.
Oversampling and the rest of the housekeeping
tanh generates harmonics; harmonics alias. The kernel runs 1×/2×/4×
(default 2×) with zero-stuffing and matched 4th-order Butterworth
anti-image/anti-alias cascades — the ladder.h pattern, self-contained here
per the house rule against shared lookup tables. Every parameter rides a
per-sample linear ramp; 16 preset slots morph through the same ramps; the
right-inlet path recomputes the coefficient per sample for signal-rate
cutoff. All state clears to zero, and all-zero state is a fixed point — a
self-oscillating patch needs a ping, exactly like the transistor ladder.
The engineering ledger
- Coupled solve vs. buffered shortcut. A "diode ladder" built from four buffered one-poles with new constants would miss the pole spread — the defining character. The 2×-larger algebra of the coupled solve is the price of the topology, paid once in closed form.
- Secant linearization vs. Newton. Newton needs the derivative of five
tanhterms through the elimination; secant gains reuse the same elimination unchanged and converge fast enough thatsolver_exactrarely iterates more than a few times. Same accuracy target, simpler code. - WDF: the documented no-go. A wave-digital rebuild of the same network
was evaluated and declined (author-approved, 2026-07-18):
solver_exactalready converges the circuit's nonlinear equations, and a WDF would re-solve the same network differing only through the Shockley-vs-tanh diode curve, with no measured reference showing an audible delta to chase. The evidence lives in the notebook's solver A/B matrix. - No
asym, nocomp. Both absences are circuit facts, not omissions: complementary diode pairs, and a high-pass that is its own passband compensation.
Checkpoint
A diffusion chain whose transfer function matches the published analysis to 0.028 dB, with the oscillation threshold derived at k = 17 and then — because the feedback high-pass is modeled rather than idealized away — never reached at stock settings, exactly like the hardware. The nonlinearity sits on the coupling edges where the circuit puts it; the coupled ZDF system is eliminated to closed form and solved once or iterated to convergence, with −44.9 dBr between the two answers at settings the hardware can't reach. The character is the coupling, and the coupling is solved, not approximated away.
The couplings are the instrument: tb303_voice.h
The field-guide chapter argued that the 303 is unmistakable because its blocks are coupled — accent reaches the filter and the amplifier through shared circuitry with memory, slide is gate behavior, the envelopes have fixed interrelations. This appendix walks the per-sample code that implements those couplings: the measured envmod law, the C13 accent-sweep capacitor, the slide one-pole, the square shaper, and the phase-2 VCA. The filter itself is the previous appendix; this file composes it.
Sources, and the division of labor between them: Open303 (Robin Schmidt) supplies the measured constants — knob travels, envelope times, the envmod mapping, the square-shaper curve; the Devil Fish documentation (Robin Whittle) supplies the circuit behavior of the envelope/accent path, including the one place this kernel deliberately diverges from Open303. Every constant in the header carries its source.
One sample, in order
process() reads top to bottom as the signal path: pitch (with slide) →
envelopes → the C13 update → the cutoff sum → oscillator and shaper →
coupling high-pass → diode ladder → VCA → output coupling. Each stanza
below is one of those steps.
The file, as a schematic. Grey is what every clone has; red is what accent touches; amber is the cutoff CV that C13 leans on.
Slide: one coefficient, no special case
m_pitch += (m_pitch_target − m_pitch) · m_slide_coef
That is the entire slide implementation: a true RC lag (τ = the slide
parameter, stock 60 ms — Open303's slideTime) on the pitch target. The
gate logic makes it behave like the hardware: note_on with the gate low
snaps m_pitch to the target before retriggering (a fresh note starts in
tune); set_pitch with the gate held moves only the target, so the lag
glides and neither envelope retriggers. Legato is slide — which is why
the sequencer's gate-hold trick (see step_seq.h) needs no slide
wire of its own.
Two envelopes, both RC discharges
The Main Envelope Generator and the VCA envelope are the same primitive — one-pole rise, exponential decay — with different constants and one coupling each:
- MEG: 3 ms attack; decay = the
decayknob (200 ms–2 s)… unless the note is accented, in which case the hardware bypasses the pot and runs at ~200 ms (accdecay, a bend, adjusts this clock). Faster and hotter is half of what "accent" means. - VCA env: fixed — ~3 ms attack (the Devil Fish "Soft Attack" bend widens it to 0.3–30 ms), a measured 1.23 s decay with no sustain, chopped by a 2 ms release at gate-off (Open303 measures ~1 ms; 2 is click-free). No knobs on the hardware, so no knobs here.
C13: the wow, as three lines of code
The accent sweep circuit is a diode feeding a capacitor through the resonance pot. The kernel's model is exactly that sentence:
drive = accent_knob · note_accent · meg
if (drive > c13) c13 += (drive − c13) · charge // diode conducts: τ = 47 ms (47k·1µF)
c13 −= c13 · drain // always draining: τ ≈ 150 ms
The diode gating (if drive > c13) is the memory: between closely spaced
accents the drain doesn't finish, so the next accent starts from residual
charge and peaks higher — the build-up. The notebook measures the cutoff
peak growing ×1.94 across a run of accents and returning within ×0.998
once they stop. The cutoff contribution combines the capacitor voltage with
a direct MEG term reduced by it (Devil Fish: "~100/147 of the MEG minus the
capacitor voltage" — what rounds the first accent's curve):
res_mix = 0.3 + 0.7·min(resonance, 1) // the pot is ganged with resonance
acc_oct = 2.0 · res_mix · (0.4·max(drive − c13, 0) + c13)
Two things to note honestly. The RC time constants are component-derived; the sweep span (2 octaves) and the 0.4 direct weight are informed approximations, flagged as such in the header. And this is the kernel's one deliberate divergence from Open303, which models its accent path as a plain 15 ms leaky integrator with no across-notes memory. The A/B was done for real — Open303 built and rendered side by side — and the Devil Fish circuit description won because the memory is documented hardware behavior. The divergence is recorded in the header, not buried.
The cutoff sum: a measured law, not a mixer
envmod is not "envelope amount into a summing node." Open303 measured the
hardware's actual mapping (calculateEnvModScalerAndOffset), and the kernel
uses those regression lines verbatim. With c the knob's log-position
between the measured travel endpoints (302…2394 Hz):
scaler = (1−c)·(3.774·e + 0.737) + c·(4.195·e + 0.864)
offset = 0.0483·c + 0.2944
fc_eff = cutoff · 2^( scaler·(meg − offset) + acc_oct )
The offset term is the hardware's "gimmick": turning envmod up also
injects a counteracting DC shift, so the sweep's resting point moves down
as its depth grows — roughly 2/3 of the sweep lands above the knob position
and 1/3 below. That interaction is why the knobs feel like a 303 rather
than like a synth with the same ranges. Note acc_oct adds outside the
envmod scaling: in the circuit the accent sweep injects directly into the
cutoff sum, so accents quack even with envmod at zero.
The square that isn't
The 303's square is its saw pushed through a transistor shaper, and Open303
measured the resulting curve. The kernel takes the polyBLEP saw
(vco.h's machinery), makes a half-cycle-shifted copy, and applies the
measured shaper:
square = −tanh( 10^(36.9/20) · shifted + 4.37 )
That ~70× gain and the 4.37 bias produce the rounded, notched pulse whose
spectrum audibly differs from an ideal 50 % square. The waveform
parameter is a ramped blend between saw and shaped square, so switching
glides click-free.
The couplings at the edges: two high-passes
Two one-pole high-passes bracket the filter — 44.5 Hz before it, 24.2 Hz
after (both Open303-calibrated coupling corners). The post-filter one earns
its keep twice: it is the output coupling, and in vca warm mode it
absorbs the saturator's signal-dependent DC, which is exactly what the
hardware's coupling capacitor does.
The phase-2 VCA: distortion that tracks the envelope
vca clean is a multiply — bit-identical to phase 1. vca warm models the
one-transistor class-A stage as a slope-normalized biased saturator applied
after the envelope gain and before the output coupling (the hardware
order):
S(v) = ( tanh(d·v + b) − tanh(b) ) / ( d·sech²(b) ), d = 2.0, b = 0.3
Unity slope at zero means quiet notes pass essentially clean; the bias
means hot signals pick up even harmonics and compression. Because the
envelope sits inside v, the distortion tracks it: measured 5.4 %
difference-signal on quiet notes, 11.5 % on full accents, ~11 % second
harmonic on a full-scale sine with ~−4 dB of compression. d and b are
probe-calibrated informed constants — the header flags schematic-derived
values as an audition-time refinement, which is the honest state of things.
Per-unit spread: seed/tolerance
The house vco.h convention, applied to a whole voice: tuning trim, cutoff
scale, envelope times, slide and C13 RCs each take a deterministic per-seed
offset scaled by tolerance, and the oscillator receives the seed plus a
proportional imperfect amount. tolerance 0 is the nominal schematic,
bit-identical to an unseeded voice (pinned by test); an mc. stack with
different seeds drifts apart the way a wall of real units does.
The engineering ledger
- One object, not a modular kit. The C13 path touches the MEG, the resonance knob, and the cutoff sum; accent touches the MEG clock, the VCA gain, and the sweep. Decomposed into osc + filter + env externals, every one of those wires would be the user's problem and most patches would omit them. The couplings live between the blocks, so the object boundary goes around them.
- Measured constants over derived ones, where measurements exist. Open303's envmod law and shaper curve are adopted verbatim rather than re-derived from the schematic — they were measured against hardware, and re-derivation would add error, not rigor. Where Open303 simplifies (the accent memory), the circuit description wins instead. Each choice is sourced at the constant.
- The wow's parameters are honest approximations. Sweep span and the direct weight await a hardware-calibration pass; the shape (diode gating, two RCs, resonance ganging) is circuit-derived and pinned by the ×1.94 measurement. Flagged, isolated, waiting — the autowah pattern.
process_at()per sample. Pitch (note + tuning + slide) can change every sample, so the oscillator is driven at signal rate rather than through a control-rate frequency parameter. The slide RC would be audibly steppy any other way.
Checkpoint
A voice whose per-sample loop is the schematic's block diagram: slide as one RC coefficient plus gate logic, envelopes as discharge curves with the hardware's fixed interrelations, accent as a hotter-and-faster MEG plus a diode-gated capacitor whose leftover charge is the wow, a cutoff law measured off real hardware complete with its gimmick, a square that is a shaped saw because that's what a 303's square is, and a VCA whose warmth tracks the envelope because the envelope sits inside the saturator. Every constant carries its source, and the one divergence from the reference implementation is documented with its reason.
One network, eight voices: the tr808_* headers
Roland built an entire drum machine out of about four circuit ideas, so the kernel does too: a bridged-T resonator class, a six-oscillator metal bank, a noise/VCA toolkit, and eight thin per-voice headers that compose them. This appendix covers the shared blocks' math — the bridged-T's trapezoidal solve and why the bass drum needs it solved that way, the metal bank's tolerance model — and the per-voice compositions, ending with the calibration pass that re-fit the family's envelopes against a real unit.
Provenance: the Werner–Abel–Smith papers (the DAFx-14 bass-drum analysis
and the cymbal/cowbell companions) and the TR-808 Service Notes, read
component by component. Every constant in these headers carries a schematic
designator or a paper section; the calibration numbers live in
tr808_calibration.ipynb.
bridged_t.h: the universal voice circuit
The network every voice reuses, with the kick's circuit-bending drawn in red.
An op-amp with a bridged-T network in its feedback path — capacitive arms
C_a, C_b, a resistive bridge, a resistive leg to ground — rings when
kicked, as a decaying pseudo-sinusoid at
fc = 1 / ( 2π · sqrt(R_leg_eff · R_bridge · C_a · C_b) )
where R_leg_eff is the leg in parallel with every resistive injection
into the center node. Roland used this network in every voice: as the
resonator of the kick, snare, toms/congas, rimshot, and claves, and as the
band-pass of the clap, cowbell, cymbal, and hats. One class, one family.
Two implementation decisions matter:
- The topology is reproduced, not summarized. With injections grounded,
the class's transfer function matches the DAFx-14 paper's printed
Eqn. (5) coefficient by coefficient (β₂ = α₂ = R_eff·R167·C41·C42, and so
on); the injected paths match their Hbt2/Hbt3, interchanged by injection
resistor; and the center node the paper calls
Vcommis exposed, because the bass drum's pitch-sigh nonlinearity reads it. The whole thing was re-derived by nodal analysis and pinned by unit test — the paper is trusted, then verified. - Trapezoidal on the states, not bilinear on the coefficients. The
discretization uses capacitor companion models — a 2×2 linear solve per
sample — which is algebraically the bilinear transform the paper uses,
but solved on the network states directly. The reason is the bass drum:
its leg resistance is modulated per sample (the attack shift shorts a
resistor through Q43; the pitch sigh shrinks the effective leg through a
fitted nonlinearity). With a coefficient-form biquad that would mean a
full redesign every sample; with the companion-model solve, a
time-varying resistor is just a changed matrix entry. Same ZDF family as
the house
svf.h.
The kick, since it exercises everything
tr808_kick.h composes the resonator with the paper's full block diagram:
pulse shaper → retrigger network → bridged-T with a feedback buffer closing
a regeneration loop → tone → level. The three signature behaviors are all
emergent from the modeled schematic: for ~6 ms the envelope saturates Q43
and the ring sits near ~129 Hz (the attack punch); as the envelope
collapses, C39/R161/D52 kick the center node again (the retrigger, so the
note doesn't step down); and leakage lifts Q43's base when the center node
swings below a diode drop — the paper's fitted memoryless nonlinearity
(α = 14.315, V₀ = −0.556, m = 1.4765e-5) converts Vcomm to a collector
current that shrinks the leg, so big early swings ring sharp and relax down
as the note decays. That is the sigh, and it is a different mechanism
from the attack jump — the paper's central untangling, preserved here. One
erratum survives in the header: the paper's Eqn. (9) as printed is garbled,
so the leg formula was re-derived from KCL at Q43's collector and matches
their stated limits.
Accent is the trigger voltage — 4–14 V on the bus, mapped from the 0..1 edge amplitude — exciting the network harder, not scaling the output. And filter states persist across triggers, so rolls interfere with the ringing tail: no machine-gun effect, by construction rather than by crossfade.
swing_vca.h: the small shared parts
The 808 shapes its percussive gains with one-transistor "swing type" VCAs
driven by RC discharges, not ADSRs. The header holds the three primitives
the noise voices share: decay_env (one-pole rise to a level, exponential
decay — retriggering re-aims the rise, no reset click), the linear
swing_vca gain (the hardware's "many high harmonics" are a flagged
refinement), and white_noise — a seeded xorshift64*, because the 808 has
exactly one noise generator feeding the snare's snappy, the clap, the
maracas, and the toms' noise layer, and because determinism-per-seed is a
house invariant: renders reproduce, tests pin, mc. instances decorrelate.
metal_bank.h: six squares and a spread
The metallic voices all draw on one bank of six Schmitt-trigger relaxation oscillators: nominal 205.3, 369.6, 304.4, 522.7 Hz plus the two trimmer-tuned at 800 and 540 (the pair the cowbell taps), duty 47.98 % per the paper's HD14584 analysis. Three modeling calls:
- Naive squares are faithful. The fundamentals sit below 1.2 kHz and the hash above them is immediately band-passed; the residual aliasing folds into the same inharmonic wash the circuit itself produces. PolyBLEP would be cost without benefit — a rare sentence in this repo, so it's documented.
- Tolerance is part of the instrument. The RC parts put any given
unit's oscillators up to ~20 % off nominal — the paper's measurement, and
the reason no two 808s' cymbals sound alike.
tolerancescales a deterministic per-seed spread of exactly that width. This is not "analog warmth" seasoning; it is a measured production statistic. - The two band-pass voicings (~3440 and ~7100 Hz, Q fit to the paper's published skirts) and the Q19 attack smoother (τ = 102.44 µs less a 0.7258 V base-emitter drop, their least-squares fit) live here too, because cymbal, hats, and cowbell all share them.
The voices, as compositions
Each tr808_*.h is a thin arrangement of the blocks above, with its own
schematic constants:
- Snare: two bridged-Ts at the late-revision ~173/336 Hz (the design
change is documented in the header), a trigger divider, and the snappy
path —
decay_env-shaped noise, band-limited near 4 kHz. - Clap (
clap|maracas): ~2 kHz dual band-pass noise through a VCA driven by the Service Notes' Figure-13 three-teeth sawtooth — the "multiple hands" transient — plus the Q70 reverberation tail. - Hats: one circuit, two envelope paths, and the hardware choke
(Q23/R173): a closed-hat trigger terminates a sounding open hat, pinned
by test. This is why
tap.808.hat~is one object with two inlets — the choke is unimplementable across separate externals. - Cymbal: the bank through both voicings with two separately enveloped bands (strike/ring/body), decay spanning the chart's 350–1200 ms.
- Cowbell: just the 540/800 pair into the ~860 Hz voicing, two-slope envelope.
- Toms/congas (
@size×@model): the resonator at the chart tunings with the D80/D81 attack pitch fall; toms add a pink-noise layer (pinned by seed-sensitivity, since the diode bend's own harmonics defeat spectral separation); congas are the same circuit, no noise. - Rim/claves: the ~1667 + 455 Hz crack with the swing-VCA's tanh harmonics, versus the pure ~2500 Hz tick.
The family also carries per-channel summing gains (k_tomc_mix,
k_cl_mix — the hardware's summing resistors into the mix bus): the
bridged-T's impulse gain grows with fc·Q, and before the balance pass the
high conga peaked at ~5.3 while other voices sat far lower. Every voice's
full-accent peak now lands in a consistent ~0.3–1.0 band, pinned by test.
The calibration pass: what measurement actually changed
The family was calibrated against a real unit (s/n 103852) recorded from the individual outs with knob positions encoded in the filenames — a 0/2.5/5/7.5/10 dial grid, 116 samples — so the comparison ran per knob cell, with identical measurements (spectral-peak fundamental, −40 dB decay, power centroid) on both sides. The result is a clean split:
- Frequencies: the schematics were right. Kick within 2.4 %, snare within 1.2 % (including the tone-max mode flip), toms/congas/cowbell/ claves within ~4 %. The kick needed no constant changed.
- Time: the recordings won. Tom, conga, cowbell, and clap tails roughly doubled; the snappy was band-limited and re-enveloped; the rimshot re-voiced low-dominant; the cymbal's decay span and brightness corrected; the closed hat's brightness residual later resolved by the hats' sizzle blend.
Each header carries its calibration note with numbers and residuals. The lesson is worth stating as a rule: schematics get you the frequencies; recordings get you the envelopes — decay behavior hides in pot tapers, electrolytic tolerances, and aging that no schematic states.
The engineering ledger
- One resonator class vs. per-voice filters. Eight voices reduce to ~4 blocks plus thin compositions only because the bridged-T class keeps the injected-path structure of the real network instead of collapsing to a generic biquad. The generality was free once the nodal analysis was done — and the kick's per-sample leg modulation required it.
- Behavioral envelope generators. The kick's EG is modeled as fast rise / ~1.1 ms release rather than as its own transistor network — the paper's own simplification, adopted with its citation. Fidelity effort went where the analysis said it matters (the leg, the retrigger, the sigh), not uniformly everywhere.
- The WDF door, left closed but unlocked. The flagged
@circuitupgrade path (wave digital, thesvf.htwo-circuit pattern) remains gated on an A/B showing an audible delta the informed model misses. The DAFx-14 paper's own finding — device nonlinearity matters less than folklore claims — suggests the gate may never open, which would itself be a documented result. - Determinism everywhere. Seeded noise and seeded tolerance mean every render, test, and calibration measurement is reproducible bit-for-bit. The calibration pass would have been guesswork without it.
Checkpoint
One network class matching the published transfer functions exactly and solved on its states so a time-varying resistor costs nothing; a metal bank whose ±20 % spread is a measurement, not a vibe; voices that are thin compositions with schematic-designated constants; and a per-knob-cell calibration pass that confirmed the frequencies, corrected the envelopes, and wrote its residuals into the headers it changed. Four ideas, eight voices, every number traceable.
Time as a function of phase: step_seq.h
The sequencer header is the smallest DSP file in the kernel and the one whose central decision does the most work per line: the engine owns no clock. It is handed a phase — a number in [0, 1) meaning "here is where we are in the pattern" — and everything else (the current step, whether this sample is a boundary, how far through the step we are) is derived from it, statelessly, every sample. This appendix explains why that one decision buys sample accuracy, polymeter, scrubbing, and drift-free multi-row lock for free, and then walks the three pieces built on it: the swing warp, the two emitters, and quantized recall.
Verification lives in two places:
tests/step_seq_test.cpp
(19 Catch2 scenarios, including a pairing test against the real
tb303_voice.h) and the executed
step_seq.ipynb.
The design of record is plans/tap.seq.md in the Max package repo.
Deriving the step, in O(1)
Ignore swing for a moment and the whole clock is one line:
k = floor( wrap(phase) · length )
Swing delays each odd-numbered step's start by swing/2 of a step, so the
start of step k is
start(k) = ( k + (k odd ? swing/2 : 0) ) / length
and the derivation gains one correction: compute the naive k, and if it is
odd but the fractional position hasn't yet reached swing/2, the sample
still belongs to the (even) step before it. Two comparisons, no search —
the boundaries are monotone, so the correction is exact.
A step entry is simply k != k_previous. That definition, rather than
"the clock ticked," is what makes the engine indifferent to how the phase
moves: run it backwards and entries still fire (pinned by test); jump it
and the landing step fires once; feed it a constant and nothing happens
after the first sample. reset() just forgets k_previous, so a transport
start fires its downbeat.
The one decision the file turns on — and polymeter falling out of it as arithmetic.
Why phase, not a pulse clock
The alternative — count incoming clock pulses — is how most step sequencers are built, and every one of them then grows a reset input, a position protocol, and a drift story. Deriving from phase dissolves all three:
- Sample accuracy is inherited from the phase source. The notebook measures trigger edges landing within one sample of the analytically computed boundaries — the one sample being float rounding at the boundary itself, not accumulated error.
- Multi-row lock is structural. Two rows fed the same ramp cannot drift, because neither owns any timing state that could drift. Mute one for an hour; it re-enters in place.
- Polymeter is arithmetic. A
length 12row againstlength 16rows off one ramp divides the same cycle differently — 12 and 16 entries per cycle, measured. The TR-808's triplet "pre-scale" falls out as a special case. - Position is explicit. Scrubbing, reversing, and jumping are the caller's choices about the ramp, not features the engine implements.
The cost is honest too: the engine cannot free-run. That is deliberate —
phasor~ (transport-locked or not) already exists, and a sequencer that
owns tempo is a sequencer that fights the transport.
Position within the step, and the gate duty
The tick also reports pos — the fraction of the current step's actual
(swung) span elapsed — computed from the same start() function. Gate
timing hangs off it: the note row closes its gate at pos ≥ 0.5, the
pinned Open303 duty. Measuring duty against the swung span rather than the
nominal step means gates never collide however hard the swing is pushed.
The trigger row: an impulse and a re-arming gap
trigger_row is the small emitter: on entry to a sounding step, emit the
step's velocity for one sample (or pulse_ms worth, for envelope
consumers), else zero. The single-sample default is a contract, not a
simplification: every downstream tap.808.* voice re-arms its edge
detector below 1e-3, and the test suite pins that two adjacent sounding
steps produce two clean detectable edges. The header documents the one way
to defeat this — a pulse_ms longer than a step merges back-to-back
triggers — rather than silently preventing it.
The note row: a five-state sentence
note_row implements the tap.303~ contract, and its entire behavior fits
in one paragraph of code. On entering step k: if the step is gated and its
slide flag is set and a note is already sounding, change the pitch
output and leave the gate level alone — that is legato, and the voice's RC
does the glide. If gated without that condition, set the gate to 1.0 (2.0
if accented) — a fresh edge. If not gated, drop the gate. Between entries:
close the gate at the duty point unless the next step is gated and
slid — that look-ahead read is the gate-hold, and it is read live from the
pattern each sample so an edit lands immediately.
Three edge cases are worth naming because the tests pin them:
- Slide from a rest is a plain trigger — there is nothing sounding to
slide from, so the flag degrades gracefully (the voice's
notemessage behaves identically). - Chained slides chain — each held boundary defers the duty close to the next step, so a run of slid steps is one unbroken gate. Sixteen gated steps with three slide flags produce exactly thirteen note-ons, measured.
- The wrap is a boundary like any other — a slide from step 15 into step 0 holds across phase 1→0, because nothing in the derivation treats the wrap specially.
One convention deserves its provenance note: the slide flag sits on the
target step (the note being slid into), matching the package's
note <pitch> [accent] [slide] message and the original interface dry-run.
The hardware stores the flag on the source note ("slide to next"). The
data models convert trivially — shift the flag column by one — and the
divergence is documented in the header rather than discovered by a user.
Quantized recall: swap on the boundary sample
Patterns live in 16 slots. recall arms rather than acts (unless
quantize now): the armed slot is applied on the next cycle entry (step 0)
or step entry, and — the detail that keeps it exact — the engine then
re-derives the current step against the new pattern's grid on that same
sample, since the new pattern may have a different length. The notebook
pins the semantics end to end: armed mid-cycle, the running pattern
finishes its bar at its own amplitudes, and the first trigger after the
wrap carries the recalled pattern's. That one message is the TR-808's
A/B-half and basic/fill switching.
What is deliberately absent
No randomness (bit-exact by construction, still pinned by test, because
invariants that aren't tested rot). No allocation after prepare() — the
pattern store is a fixed 64-step array times 16 slots. No run/stop, no
direction modes, no ratchets: the first two belong to the phase source, and
the last is a future emitter, which is the point of the next paragraph.
The engineering ledger
- Engine/emitter split. The clock math lives once;
trigger_rowandnote_roware each a screenful. A future row flavor — CV, probability, ratchet — is another emitter, not another clock. This is also why the Max-side question "one generic object or two family objects?" could be answered by product taste rather than by implementation cost. - Look-ahead vs. cached hold. The gate-hold could cache "next step slides" at entry; reading it live costs one array access per sample and makes pattern edits take effect mid-step. Cheap beats stale.
- Sample-resolution boundaries. Sub-sample trigger placement (fractional
edge amplitudes à la BLEP) was considered and declined: the consuming
voices detect edges at sample resolution, so sub-sample machinery would
add complexity no consumer can observe. If a future voice interpolates
its trigger time, the
tickalready carries the information needed to add it. - The armed-recall re-derivation. The subtle bug in naive quantized recall is applying the swap after deriving the step, leaving one sample computed against the old grid. Applying, then re-deriving within the same call, is two extra lines and the difference between "exact on the wrap sample" (measured) and "usually fine."
Checkpoint
A sequencer that is a pure function of phase plus a pattern: one line of derivation, one comparison for swing, entry as inequality — and from that, sample accuracy, polymeter, reversibility, and drift-free lock without a clock to maintain. The rows translate steps into the two shipped voice contracts, with slide as a held gate and a live look-ahead; recall swaps patterns on the exact boundary sample. Nineteen scenarios and an executed notebook agree, and the most satisfying number in either is small: thirteen note-ons, for sixteen steps, three of which arrived without knocking.
Three ways to move a pitch: yin.h, psola.h, pvoc.h
The user-facing chapter promised that three interchangeable engines land the same intonation. This appendix is about why that is hard: each engine is a claim about what a pitched sound is, and each claim fails somewhere specific and measurable. Two of those failures were found the good way — as failing tests during development — and both are now pinned as contracts rather than patched into vagueness.
These three headers live in the shared DspTap repository (the same home
as the real FFT that machine/spectral.md describes), because a pitch
detector and two shifters are not Max material or even TapTools material —
they are primitives, in the fft.h mold: a double-precision golden model,
a float32 embedded profile pinned against it, allocation-free noexcept
processing, fixed documented latency, and hot loops kept contiguous as
future Helium/HVX backend seams. Every number below is produced by the
shipping code through DspTap's C ABI in notebooks/pitchshift.ipynb, and
gated in test_yin.cpp / test_psola.cpp / test_pvoc.cpp.
A period is a lag that explains the signal: yin.h
Autocorrelation says: a signal is periodic at the lag where it best matches itself. The trouble is that a harmonic-rich signal matches itself rather well at twice the true period too, and "rather well" wins often enough to make naive autocorrelation an octave gambler. YIN (de Cheveigné & Kawahara, 2002) replaces "best match" with "smallest failure": a squared difference function
d(τ) = Σ (x[j] − x[j+τ])², j over the integration window
then divides each lag's failure by the running mean of all failures up to
that lag — the cumulative-mean normalization — so d′(0) ≡ 1 and small
lags stop being free wins. The detector takes the first lag whose
normalized failure dips under an absolute threshold (0.1 by default),
descends to the local minimum, and refines it with a parabolic fit over the
three surrounding values — the sub-sample step that turns an integer lag
grid into a fractional period.
The contract, measured: worst sine error 0.17 cents across 82–988 Hz (including deliberately non-integer periods), worst sawtooth error 0.155 cents with no octave errors — the trap the normalization exists to disarm. Noise and silence report unvoiced, and the threshold gates honestly (a deliberately dirtied sine flips to unvoiced when the threshold is tightened below its measured aperiodicity).
One honest limit, kept on purpose: first dip under threshold scans from short lags to long, so on synthetic material whose fundamental is nearly absent — a formant bump with almost no energy at f₀ — a subharmonic lag that happens to land on an exact integer can dip deeper than the true period's slightly-off-grid dip, and the detector follows it. The notebook demonstrates this deliberately and measures such material with a cepstral oracle instead. Real voices keep enough fundamental that the rule holds; the failure is documented, not hidden.
Finding 1 — a shifter that moves everything except the envelope: psola.h
TD-PSOLA's move is disarmingly physical. Put an analysis mark every period. Cut a two-period Hann grain around each mark. To synthesize a new pitch, lay the grains back down at a new spacing — period/ratio — and sum. The windows are arranged to sum to one at the identity, grains are scaled by 1/ratio to keep the sum flat elsewhere, and synthesis marks are placed with sub-sample precision (each grain resampled through the same 4-point Hermite kernel the rest of the family uses) so mark rounding never becomes pitch jitter. The file adds one real-time honesty: marks come from a free-running period-synchronous scheduler, not glottal-epoch estimation — the standard practical simplification — and the caller supplies the period, so detector and shifter stay independently testable.
Then the finding, told as it happened. The first shift-accuracy test fed the shifter a pure sine at ratio 2 and got back 0.0000 — silence, from a correct implementation. Because that is what PSOLA does: re-spacing period-synchronous grains resamples the source's spectral envelope at the new harmonic spacing. A voice's envelope is wide — formants — so the new harmonics sample it fine, which is exactly the celebrated property: formants stay put while pitch moves. A pure sine's envelope is a single spike at f₀, and after an octave up the new harmonic grid (2f₀, 4f₀, …) contains nothing at f₀ — the output honestly, correctly vanishes. One property, two faces.
The response was not to patch the algorithm into something less itself. The
shift tests were rewritten onto voice-like material (a normalized
band-limited sawtooth: ±8 cents across ratios 0.5–2.0 at healthy level),
and the pure-tone behavior got its own pinning test,
PureToneOctaveUpThinsOut, so that if this property ever changes, someone
is forced to explain why. The header now opens with the warning label:
know what PSOLA is; feed it harmonics; for pure tones use a
waveform-preserving shifter.
Latency is fixed at 2·max_period + 2 samples — the price of grains that
must be fully received before they can be laid back down.
Finding 2 — the naive phase vocoder loses half its level: pvoc.h
The textbook pitch shifter looks like four honest lines: STFT with Hann
windows at 4× overlap; per-bin instantaneous frequency from the
frame-to-frame phase increment; remap each analysis bin k to synthesis
bin round(k·ratio); accumulate each synthesis bin's phase at its scaled
frequency and inverse-transform. It is in tutorials everywhere. Measured on
a unit sine, it delivers 0.14–0.46 of the input level at fractional
ratios — more than half the signal simply gone — and its "identity" at
ratio 1 is a sine of the right frequency with the wrong waveform.
Two structural reasons. First, a single partial does not live in one bin;
it lives in a Hann mainlobe pattern across four-ish bins, and
round(k·ratio) scatters that pattern (220 Hz × 1.5 lands on bins
{5, 6, 8, 9} — nothing at the true target, 7.04). Second, free-running
per-bin phase accumulators destroy the phase relationships across the lobe,
and the overlap-add — which is a resampling filter with real opinions —
partially cancels what remains.
The shipping design is Laroche–Dolson peak-region shifting. Find the spectral peaks (local maxima over ±2 bins, gated 80 dB below the frame's strongest bin so the noise floor cannot claim regions). Split the spectrum into regions around them. Translate each region rigidly by an integer bin offset — the lobe pattern survives intact, phase relationships and all — and rotate the whole region by a single accumulated per-hop phase ψ.
And here is the bug that cost an afternoon and earned its own comment
block: ψ must accumulate the full per-hop frequency difference,
ψ += 2π·hop·f·(r−1)/N, not the sub-bin residual left after the integer
shift. An integer bin shift is implemented, in effect, by a modulator
e^(2πi·shift·n/N) — but n is the frame-relative sample index, so that
modulator restarts every frame and contributes nothing to frame-to-frame
phase advance. Subtract the shift from ψ (the "obvious" refinement) and
every frame disagrees with the last about where the shifted partial's phase
should be; the overlap-add quietly shreds the signal. With ψ carrying the
full difference, the measured contracts land: sub-cent frequency accuracy
at every tested ratio, ~0.95 level everywhere, and — because at ratio 1
every shift and every ψ increment is exactly zero and the analysis phases
pass straight through — exact waveform identity, one frame late, to
7.8 × 10⁻¹⁶.
The envelope as a filter: LPC formant preservation
set_formant(true) adds the classic source-filter correction. Per analysis
frame: autocorrelate the windowed time frame to lag 48, run Levinson–Durbin
(always in double — an order-48 recursion in float32 is not a place to
economize), and evaluate the prediction polynomial's magnitude over all
bins with one extra FFT of its coefficients — the same transform engine,
one more call. That gives a spectral envelope E(k) = 1/|A(e^jωk)|, and
every relocated bin trades envelopes: content moving from bin k to bin
j is scaled by E(j)/E(k) (clamped to ±24 dB so a near-zero envelope
cannot mint gain). The excitation moves; the envelope stays. Measured: a
synthetic 800 Hz formant on a shifted-up-a-fifth voice stays at 800 Hz
with the flag on (band-energy ratio 62:1) and dutifully chipmunks to
1200 Hz with it off. At ratio 1 the correction is E(k)/E(k) — exactly
unity — so the identity contract survives the feature untouched. The
method is implemented from the published literature only, which in this
corner of DSP is a policy statement, not just a citation habit.
The house pattern
All three files repeat the fft.h discipline because it keeps paying:
basic_*<Sample> templates with double as the golden model and float
as the embedded profile, cross-precision agreement pinned by tests;
geometry fixed at construction and every buffer allocated there;
noexcept, allocation-free processing; latency as a number in the header,
not a vibe; and the expensive inner loops (YIN's difference function above
all) written as plain contiguous arithmetic so a Helium or HVX backend can
slot in behind the same contract with the scalar build remaining the
oracle.
Checkpoint
A detector that measures failure-to-match instead of match, normalized so short lags stop cheating, refined below the sample grid — sub-cent, octave- safe, honest about the one synthetic that fools its first-dip rule. A grain shifter whose deepest property — resampling the spectral envelope — is both its celebrated feature and its pure-tone failure, pinned from both faces. A phase vocoder that works because peaks move as rigid families with one phase register each, carrying the full frequency difference — since the integer shift's modulator restarts with every frame. And an LPC envelope trade that lets the excitation move while the mouth stays. Three claims about what a pitched sound is; three sets of receipts.
The nearest allowed note: tune.h
The pitch-primitives appendix built the parts: a detector and
two shifters, each with a numeric contract. This appendix is about the
composition — tap::tools::tune::corrector, the object behind
tap.tune~ — where the interesting problems are not algorithms but
policies: what runs when, what is allowed to allocate, what happens when
the detector reports nothing, and one measured surprise that became the
kernel's best war story. The scenarios in tune_test.cpp drive the class
directly, using the DspTap detector as an independent pitch oracle on the
output; notebooks/tune.ipynb re-measures the headline claims through the
C ABI.
The pipeline and its clock
process() runs per sample; analysis runs per hop (256 samples at 48 kHz,
about 5.3 ms, scaled with the rate). Each sample: feed the detector's input
ring, maybe analyze, advance two slews (the applied correction and the
grain window), compute the ratio, resynthesize.
The geometry trick that keeps every setter real-time safe: the detector is
built at prepare() for the worst case — lags from 2 kHz down to 55 Hz —
and the user's set_range() merely filters results afterward, treating
out-of-range estimates as unvoiced. Changing the range never reallocates,
so it is safe mid-audio, and the price is a fixed analysis cost: with
window = τ_max = 873 samples at 48 kHz, the YIN difference function is
roughly 760k multiply-adds per analysis — an ~80 µs scalar spike every
5.3 ms, well inside a 64-sample vector's 1.3 ms budget, and the
FFT-accelerated difference function remains available behind the same
contract if an embedded target ever objects.
Analysis converts the detected period to MIDI, chooses a target (next section), sets the correction goal in semitones, and retargets the grain window. Unpitched frames set the correction goal to zero — the corrector relaxes toward honesty — while the window holds its last value rather than lurching toward a default.
The pipeline and its two clocks — the dashed region runs per hop, everything else per sample.
The period lock, told as a bug hunt
The first version of the resynthesis stage was the two-tap tap.shift~
engine with its grain window clamped to a sensible fixed range, minimum
5 ms. The oracle tests immediately failed — not wildly, musically: a
452 Hz input hard-snapped to A440 came out at 441.4 Hz, 5.4 cents sharp.
Detection was exonerated first (0.06 cents), then the applied correction
(right to five decimals). The bias lived in the shifter itself, and an
isolation experiment found the shape of it:
| grain window | measured output | error |
|---|---|---|
| exactly 2 detected periods (212.4 smp) | 439.98 Hz | −0.06 cents |
| fixed clamp (240 smp) | 441.38 Hz | +5.41 cents |
| 480 smp (≈4.52 periods) | 441.36 Hz | +5.34 cents |
The two taps ride the same phasor half a cycle apart, so they sit window/2 samples apart in the delay line. When window/2 is an integer number of source periods, the taps read the same phase of the waveform and their crossfade is invisible — and the average retune ratio is exactly the phasor's ratio. When it isn't, every crossfade splices a phase jump into the output, and the jumps do not average away: they bias the pitch. Period-locking the window is not a quality nicety; it is what makes the ratio true.
The fix: the window targets the smallest even multiple of the detected period that clears the minimum — 2 periods normally, 4 for high pitches whose 2 periods would be under 5 ms — so the taps always sit an integer number of periods apart. The clamp survives only as an outer bound. This is the cleanest example in the book of a defect no assert-on-internals test would ever catch: only an oracle — the detector listening to the output — could hear 5 cents.
Choosing the target
Scale mode: round the detected MIDI to a center note, then scan offsets
−6…+6 for enabled pitch classes, keeping the candidate nearest to the
fractional detected pitch (ties resolve to the smaller motion). A tritone
of search radius suffices for any non-empty mask; an empty mask returns no
target, and no target means a zero correction goal — the object never
guesses. MIDI mode is simpler and blunter: the nearest currently-held note
in absolute MIDI space, whatever the distance (clamped to ±12 semitones of
actual correction). amount scales the goal before the glide — a fader on
the distance itself.
The glide
The applied correction chases its goal through a one-pole with time
constant speed (0 = assignment, the hard snap), evaluated per sample so
the goal can move every hop while the glide stays silky. The ratio is then
2^(applied/12), computed per sample; the exp2 is cheap and the
alternative — caching with edge cases — is not. The grain window rides its
own 15 ms slew toward the period-locked target, and the detected period
gets a third slew for the PSOLA backend's per-sample period input. Three
small slews, no zippers, no special cases at the joins.
Three backends behind one seam
set_backend() swaps only the last stage; detector, mapper, and glide are
shared state that survives the switch. Both alternate engines are
constructed at prepare() (PSOLA sized to the deepest detectable period,
the phase vocoder's FFT scaled to ~21 ms at any rate), so switching
allocates nothing and is safe mid-audio; the incoming engine is cleared to
silence first — a fade-in, not a splice of stale buffers. The ledger, at
48 kHz:
| backend | resynthesis | latency |
|---|---|---|
| grain | period-locked two-tap (in-kernel) | ≈ base delay + window/2, a few ms |
| psola | tap::dsp::psola | 2 × 873 + 2 = 1748 samples ≈ 36 ms |
| pvoc | tap::dsp::pvoc (+ optional LPC formant trade) | 1024 samples ≈ 21 ms |
The backend-parametrized scenarios feed all three the same 46-cent-sharp
sawtooth and require the same landing (±6 cents at healthy level); a
switching scenario hops between engines mid-signal and requires finite
output and a correction that is still standing at the end. set_formant()
forwards to the phase vocoder — PSOLA preserves formants by construction
and the grain engine is waveform-preserving, so the flag deliberately
touches one path.
Learning the key
The auto-key learner is thirteen doubles and a policy. Every voiced analysis adds 1 to its pitch class's histogram bin; every analysis multiplies all twelve bins by a leak chosen so the histogram forgets with a 60-second time constant. On demand — never on a schedule — the histogram is scored by Pearson correlation against the published Krumhansl–Kessler major and minor profiles at all twelve rotations, and the best of the 24 becomes the estimate, with the winning correlation as confidence. A mass guard withholds any estimate until roughly half a second of voiced material exists, so silence cannot have an opinion.
The design decision that matters is that the learner is advisory:
autokey_estimate() reports and autokey_apply() adopts, but nothing in
the audio path ever re-aims the targets on its own. This is a UI-safety
argument, not modesty — a corrector that changes its own scale mid-phrase
turns a wrong estimate into a wrong performance, and the person at the
patch cannot undo what they never saw happen. Measured: a tonic-weighted
D-major scale scores D major at 0.95 confidence; an A harmonic-minor
melody scores A minor; reset withdraws the estimate.
The ledger
- Allocation discipline. Everything sized at
prepare(): detector frame and ring, both alternate backends, the grain buffer at the maximum window. After that the audio path allocates nothing; every setter either writes a double, flips a flag, or clears preallocated state. - Oracle-based testing. The scenarios measure the output's pitch with the independently-certified DspTap detector — the only kind of test that caught the period-lock bias — and use sawtooth, not sine, wherever PSOLA participates, per its documented material contract.
- The maxtest. One assertion runs inside a real Max: unpitched DC in, exactly DC out — detector unvoiced, correction zero, complementary envelopes summing to one. It pins the whole "never guess" policy at unity gain in the shipping binary.
- What the wrapper adds. Only plumbing: attribute forwarding, the
atomic pitch handoff to a scheduler timer for the right outlet, and
applykeywriting back through the attributes so Max's saved state stays the source of truth.
Checkpoint
A per-sample corrector with a per-hop brain: worst-case geometry bought at
prepare() so nothing ever allocates again, unpitched input relaxing to
zero correction, and three slews smoothing every join. The war story is the
period lock — two taps half a window apart are only honest when that
half-window is an integer number of periods, and only an oracle test could
hear the 5-cent lie. Targets are chosen, never invented; the glide is one
pole and one exp2; the backends swap behind a seam that clears to silence;
and the key learner watches, scores, remembers for a minute — and speaks
only when spoken to.
The clipper in the loop: overdrive.h
The user-facing chapter claimed that tap.overdrive~'s gain tilts with
frequency and that the tilt grows with drive — behavior a memoryless
waveshaper cannot produce. This appendix derives the loop that produces it,
shows why the obvious implementation of that loop is a stability bomb and how
the file defuses it, and records the design decisions — shaper choice,
asymmetry mechanics, oversampling versus ADAA — with the alternatives they
beat.
The design brief was not a schematic (none is published for the Little Green Wonder, the listening reference): it was the class of TS-lineage feedback overdrives. The honest statement of the goal, from the project's handoff notes: the interesting part is not the transfer curve — it's the frequency-dependent gain and the softer, never-fully-flat knee that a feedback clipper gives you.
The topology, and what it must do
In a TS-lineage pedal the diodes sit in the feedback path of a non-inverting op-amp stage whose feedback network is frequency-dependent. Two consequences:
- The loop gain — and with it the effective clip threshold — varies with frequency: bass sees little gain and stays clean, mids see all of it.
- The output is
input + limited feedback term: even at maximum drive the transfer's slope never reaches zero, because the clean input always passes.
overdrive.h models this with the minimal structure that keeps both traits:
w = shape( G·x − g_fb·LP(w) ) the clipper inside a lowpass feedback loop
y = x + w the unity clean path (non-inverting topology)
The whole kernel on one line. The red loop is the frequency-dependent gain; the amber path is why the transfer never flattens; everything inside the dashed region runs at the oversampled rate.
G is the drive gain (a dB sweep, +6 to +46). The lowpass LP (one-pole,
corner 660 Hz) makes the fed-back signal predominantly low-frequency, so the
negative feedback suppresses gain exactly where the pedal does. In the linear
region (shape ≈ identity) the loop's small-signal gain is
w/x = G / (1 + g_fb·|LP(ω)|)
— at DC, G / (1 + g_fb); far above the corner, G. The file picks g_fb
from a single voicing constant: g_fb = G/k_lf_gain − 1 with
k_lf_gain = 2, which pins the low-frequency gain at +6 dB regardless of
drive while the mids ride G all the way up. That one line is the measured
headline — a bass-to-mid tilt of +5/+16.3/+17.2 dB at drive 0/0.5/0.9 —
and it is the real-pedal behavior: turning up a TS makes the mids filthier
while the low E barely moves.
Why the loop must be solved zero-delay
The naive discretization feeds back yesterday's lowpass state:
s = G·x − g_fb·lp_state // uses the previous sample's state
w = shape(s)
lp_state += a·(w − lp_state)
That inserts a unit delay into a feedback loop — the same mistake as the
Chamberlin SVF, with the same fuse. Linearize it: the state-to-state map has
Jacobian J = (1 − a) − a·g_fb·shape′. At 48 kHz × 4 oversampling, a 660 Hz
one-pole has a ≈ 0.021; at drive 0.9, g_fb ≈ 62. With shape′ = 1
(small signal — the quiet case!) J ≈ 0.979 − 1.34 = −0.36: stable, fine.
But push g_fb higher — drive 1.0 gives G = 200, g_fb = 99 — and
J ≈ 0.979 − 2.12 = −1.14. |J| > 1: the loop limit-cycles near Nyquist,
audible as a parasitic whine that comes and goes with the signal level.
A feedback clipper that oscillates when you turn it up is not a pedal, it's a
bug report.
The fix is the house zero-delay move (svf.h's driven circuit, ladder.h's
solver_fast): integrate the one-pole trapezoidally (TPT), solve the loop's
linear part implicitly, then apply the nonlinearity and commit its output
to the state. With the TPT one-pole v = (g·w + s)/(1 + g), substitute into
the loop and solve for the node as if shape were identity:
w_lin = ( G·x − g_fb·s/(1+g) ) / ( 1 + g_fb·g/(1+g) )
w = shape(w_lin + bias) − shape(bias)
v = (g·w + s)/(1+g); s ← 2v − s
No delay in the linear loop, so no delay-induced instability at any g_fb;
and because shape′ ≤ 1 everywhere, the committed value only ever reduces
the effective loop gain below the linear prediction — the approximation errs
toward stability. At DC the solve gives w = G·x/(1 + g_fb) exactly, which
is what makes the pinned-bass-gain arithmetic above exact rather than
approximate. The kernel suite pins the consequence: after a full-drive,
full-asymmetry signal stops, the output decays below 10⁻⁶ — no limit cycles.
The shaper: u/√(1+u²), and why not tanh
Three candidates from the brief, in the order they were rejected:
std::tanh— the reference softclip, and the expensive outlier: a transcendental call per (oversampled) sample that vectorizes badly.- Padé-style tanh approximations — cheap, but the usual forms are exact only on a bounded interval and go flat (or worse, retreat) beyond it — reintroducing the hard plateau this design exists to avoid, with a curvature discontinuity at the seam that aliases.
shape(u) = u/√(1+u²)— chosen: C∞ (no curvature seam to alias), strictly monotonic, asymptotic to ±1 but never flat, one multiply-add and one square root — which vectorizes as a reciprocal-sqrt instruction on every SIMD ISA this kernel targets, and reduces to a small LUT for a future fixed-point port.
The chosen curve reaches its asymptote more slowly than tanh — softer knee, lower-order harmonics — and unlike the hard clip it has no corner for the spectrum to pay for.
Asymmetry — the even-harmonic control the odd-only Jamoma curves structurally
lacked — is a bias inside the shaper, output-corrected so silence stays
silence: w = shape(u + b) − shape(b) with b = 0.5·asymmetry. At
asymmetry 0 the whole path is an odd function and the measured H2 sits at
the numerical floor (−151 dB); at 0.6 it is −26 dB and musically present.
The correction term keeps the first-order DC out, but a biased clipper still
rectifies: under signal it makes DC, and the feedback one-pole would
happily integrate it. Hence the DC blocker after the clipper —
y[n] = x[n] − x[n−1] + 0.9997·y[n−1], the Jamoma TTDCBlock constant kept
for provenance — permanently in the path, not an option. (The original
TTOverdrive instantiated that same blocker and then overwrote its output
buffer without using it; the vestigial call was one of the tells, noted in
the handoff brief, that the old code path was never going to be the base.)
Oversampling, not ADAA (for now)
Clipping generates harmonics without limit; everything past Nyquist folds
back inharmonically. Two published remedies: oversample the nonlinearity, or
antiderivative anti-aliasing (Parker et al., DAFx-16). ADAA is cheaper per
dB of alias suppression, but its x[n] ≈ x[n−1] fallback branch is hostile
to the branchless-SIMD constraint this kernel inherits from its embedded
targets, and its difference quotient loses precision in single-precision
float — a real concern for the fixed-point/f32 ports. So v1 oversamples:
zero-stuff + 4th-order Butterworth anti-image up, matching anti-alias down —
the ladder.h/svf.h resampler verbatim, self-contained per house rule.
Factors 1/2/4/8, default 4×. Measured on a hard-driven 5 kHz tone: the
folded seventh harmonic improves from −22 dB (1×) to −36 dB (4×) while the
in-band harmonics stay within measurement error. ADAA remains the flagged
experiment for after the voicing locks, so the comparison is apples to
apples.
Every red peak standing above the blue mass is inharmonic fold-back the default 4× removes; the true harmonics (multiples of 5001 Hz) coincide in both traces. The dashed line marks the folded seventh harmonic the kernel suite pins.
The voicing layer, honestly labeled
Everything above is structure; the sound of the body control is a handful
of constants (k_voice_* at the top of the file): the pre-clipper highpass
corner sliding 40→320 Hz across the knob, the upper-mid bell at 1150 Hz
(above the classic TS hump — the LGW's push sits higher), the +2.5 dB
counterclockwise treble shelf, the fixed +1.5 dB mid seasoning. They produce
the measured control shape (±10 dB at 100 Hz between extremes, +4 dB at the
bell) and they are by-ear placeholders: the header says so, this book
says so, and the numbers will move when the in-Max voicing pass against LGW
demos happens. What will not move is where they live — all linear EQ outside
the nonlinearity, because in the reference pedal that is what the Body knob
is.
The parameter block is normalized on purpose
drive and asymmetry are 0..1, body is −1..+1; only preamp/output
carry units (dB). The perceptual mapping (dB sweep of G, level
compensation) lives inside the kernel, not in the knob range — so the
parameters map directly to controllers, to live.dial, and to Q15/Q31
fixed-point registers on the Cortex-M targets this library's headers are
written to reach. Parameters ride the standard per-sample linear ramps
(default 20 ms); the derived coefficients — G, g_fb, the solve constants,
the voicing biquads — refresh only on samples where a ramp actually moved,
the same two-tier scheme as svf.h.
Everything in this chapter is executable: the loop math and stability claims
are pinned by tests/overdrive_test.cpp (silence decay, tilt-grows-with-
drive, even-harmonic emergence, DC blocking, alias improvement, determinism),
every number is a cell in
the verification notebook,
and the measured figures are regenerated from the shipping kernel by
book/figures/overdrive.py.
Wear as the stabilizer: tape_loop.h and discreet.h
Every regenerating loop in this library before these files made the same
promise the same way: the loop is strictly contractive because feedback is
capped below one (delay.h's k_fb_max = 0.99, the comb bank's calibrated
ring time). tape_loop.h and discreet.h exist to make the opposite
promise — regeneration at exactly 1.0, bounded anyway — and this appendix is
the derivation of why that is allowed.
A shared header, by the house rule
The family needed the same four pieces twice (discreet.h and airport.h
are both tape machines), and the reuse rule sorted them cleanly. Classes
with state went into a shared header the way swing_vca.h was created for
the drum family: tape::reel, tape::wow_flutter, tape::wear, and a
tape::ramp that is a cited copy of delay.h's anti-zipper unit. Few-line
expressions stayed copies-with-citation, as ever: the Hermite polynomial
inside reel is the same read as delay.h, line for line, and says so; the
saturator is not copied at all but included — vca::swing_shape, the shared
swing-type stage, with the reason on the include line.
reel: one wrap, two topologies
A reel is position-addressed circular storage whose reads and writes wrap
modulo a settable loop length, not the buffer size. That one decision lets
the same class serve both kernels. discreet.h runs it as a delay line:
loop length equals capacity, an integer write head advances forever (wrapped
into range each sample — a bare long head would overflow LLP64's 32-bit
long in half a day of audio), and the play head trails it by the loop
span. airport.h runs it as a true loop: length set per piece, one
free-running head, positions handed in raw because the reel does all modular
arithmetic itself. A length change is deliberately a splice — content
kept, positions re-wrapped — because that is what cutting tape does.
wow_flutter: periodic on purpose
The transport error is two sines — slow-deep wow, fast-shallow flutter —
returning a read-position offset in samples, phases zeroed at prepare().
The periodic term is the dominant one in the tape-echo literature
(Arnardóttir, Abel, Smith, AES 2008), but the deeper reason the stochastic
term is a documented non-goal is testability: the wow promise is pinned by
predicting peak pitch deviation in closed form (depth · 2π · rate, so 2 ms
at 0.5 Hz ⇒ ±10.9 cents) and measuring it with the YIN oracle — 10.9
measured — and that oracle test only exists because two renders are
bit-identical. Determinism was a design force here, not an afterthought.
wear: the boundedness argument
One pass of generation loss is three stages in fixed order: an exact
one-pole darkening lowpass (1 − e^(−2πf_c/sr), the grm_comb.h map), the
shared saturator swing_shape(v, d) = tanh(d·v)/d, and the normalized DC
blocker. Each carries one clause of the proof:
tanhis bounded, so for any drived > 0the wear output can never exceed1/d— whatever the loop has accumulated. That is BIBO stability at regen 1.0, unconditionally, from the saturator alone.- The DC blocker (pole 0.999, peak gain normalized to exactly 1 — the normalization grm_comb.h earned the hard way, chasing a +0.2 dB/s swell) kills the one frequency the lowpass would happily sustain forever with an offset attached.
- The lowpass is strictly contractive above its corner and asymptotically transparent below it — which is not a leak in the proof but the musical contract: at drive 0 and regen 1.0 the sub-corner band sustains indefinitely, cleanly. The header calls this the Frippertronics contract and states it rather than hiding it.
So where delay.h proves stability by gain, this family proves it by
shape: each pass survives because it is degraded. The pinned test drives
regen 1.0 for ten seconds of ring and asserts non-growth — never decay,
because decay would betray the contract just as surely as growth.
The doppler decision
discreet::machine gives loop_seconds an ordinary ramp and does nothing
else, because nothing else is needed: moving a fractional read head is
tape-speed doppler. A 0.5 → 0.75 s glide over half a second reads back an
octave down mid-move (measured: 220 Hz, then re-lock within five cents) with
no discontinuity, since position is continuous even where its slope is not.
The rejected alternative — crossfading between two taps — would have hidden
the machine, and hiding the machine is the one thing this kernel is for.
The wow offset is clamped so the read can never cross the record head; at
absurd depths on short loops the transport flattens against the clamp
rather than wrapping, which the header files under honest limits.
A finding: the arithmetic agreed
The per-pass wear transfer is fully analytic — regen · |H_lp| · |H_dc| on
the unit circle — so the notebook measured it the direct way: a two-tone
burst (300 Hz under the corner, 6 kHz over it) recirculated at drive 0, each
generation's tones read by Goertzel. Measured per-pass ratios: 0.292 and
0.890. Predicted: 0.292 and 0.890. Three decimals of agreement between a
rendering kernel and a formula derived independently in the test is the
cheapest kind of confidence this library knows how to buy, and both the test
(with 15% and 5% tolerance bands it never needs) and the executed notebook
carry the measurement.
The engineering ledger
The suite leans on four instruments. Analytic transfers wherever the path is linear (the per-pass darkening scenario asserts against the exact formula, both tones, both directions — highs die faster and lows barely fade, so the test cannot pass vacuously). Two-window RMS for long-run claims, inherited from the comb bank's swell story: regen 1.0 rings ten seconds and the late window may not exceed the early one. The YIN oracle for anything with a pitch: wow depth in cents against the closed form, the doppler glide and its re-lock. And bitwise assertions where the law is exact: mix endpoints, the first echo returning as literally the recorded impulse, two wow renders identical to the bit. The DC-step scenario checks the blocker's actual job — a held offset at regen 1.0 does not accumulate and the tail's mean returns below 0.02 — rather than a decay the contract never promised.
Checkpoint
One shared header, four blocks: a reel that wraps at the loop, a transport
that is two deterministic sines, a wear stage whose tanh bound is the
stability proof, and a cited copy of the house ramp. discreet.h composes
them into the two-machine loop where regeneration legally reaches 1.0,
loop moves are doppler because read heads are physical, and every claim is
carried twice — discreet.ipynb executed, discreet_test.cpp pinned.
Free-running heads, one shared clock: airport.h
airport::loop_bank is structurally the smallest kernel in the family — a
fixed array of loops, a stereo sum, no feedback anywhere — and that is what
makes it interesting to read: nearly every promise it makes is structural,
so nearly every test on it is bitwise. This appendix walks the file in code
order and dwells on the one discipline that defines it.
loop_state: the multitap idiom with a reel in each seat
The bank is std::array<loop_state, k_max_loops> with an active count —
delay.h's multitap shape, kept deliberately: per-index setters that
silently no-op on a bad index, getters that return safe defaults, newly
activated slots arriving at their stored settings. Each seat holds a
tape::reel (its own worst-case buy — eight 30-second reels is ~92 MB of
double tape, the family's largest allocation, stated in the header rather
than discovered in production), a tape::wear used as a playback shade, a
phase, a record flag, and three ramps (level, pan, darken).
The phase discipline
The load-bearing sentence in the header is "the phase is NEVER reset":
recording starts wherever the head is, set_loops activates a loop with
its head wherever it last was, a splice re-wraps the head modulo the new
length without rewinding, and only prepare()/clear() — DSP restarts —
may rewind. The reason is musical: in "2/1" the free-run is the piece,
and any convenience reset (snap to zero on record, realign on length
change) would quietly delete the composition. The pinned scenario earns the
promise the blunt way: it fires a setter storm mid-render — level, darken,
record, length, count — and then requires the click grid unmoved and the
head advanced by exactly the samples processed. phase() exists as
introspection precisely so that test could be written.
Record semantics
record is a gate, not an action: while on, the input replaces the tape
at the integer head position, after the read — so you hear the previous
generation under the head while punching, and one Hermite support point
(two samples) of the old generation blends across the punch, which the
header files under honest limits instead of papering over with a crossfade.
No overdub-sum, because the provenance had none: each Airports phrase was
recorded once. Freeze is the strong promise — record off, and two
successive passes of the loop are required bit-identical. That promise is
only possible because of the next decision.
The shade and its bypass
Per-loop darken reuses tape::wear with drive pinned at 0, as a static
playback tone — deliberately not generation loss, because a frozen loop
replays the same magnetic imprint every revolution and modeling wear on it
would be dishonest physics. At the band ceiling (the default) the stage is
bypassed entirely: not "flat enough", but not-in-the-signal-path, which is
what upgrades the freeze test and the hard-pan test (a pan of −1 adds the
loop's samples to the left bus unscaled) from tolerance checks to bitwise
facts. Engaged, the shade is the exact one-pole from grm_comb.h, and the
notebook measures a 6 kHz phrase through a 1 kHz shade at 0.169 of its
transparent twin against 0.169 predicted.
composite_period_seconds
The lcm of the active loop lengths in samples, folded pairwise with a
long long gcd, overflow detected before each multiply and reported as
+inf. It is introspection, not DSP — but it is the piece's thesis as a
number: 24000- and 30000-sample loops report exactly 2.5 s (and the pinned
scenario also proves the rendered output repeats at 120000 samples and
does not repeat at 60000), while seven airport-scale lengths overflow to
infinity, which the header calls the point.
A finding: the raster before the assertion
The lcm scenario existed as an assertion first — bitwise equality of two 2.5-second windows — and it passed, which is exactly why it was worth plotting. The notebook's event raster (every return of loop A, loop B, and their sum on one timeline) made the same fact visible: the coincidence pattern audibly and graphically re-enters at 2.5 s and drifts everywhere short of it. The assertion pins the promise; the raster is what convinces a human the promise means something. The pair — one bitwise test, one executed figure — is this library's preferred way to hold a structural claim from both sides.
The engineering ledger
Almost everything here is exact, so the suite asserts exactly: bit-equality
for freeze and for the lcm window, bitwise silence on the far bus for hard
pans, phase() continuity to 1e−9 through the setter storm, and the splice
law (0.9 of a 1 s loop re-wraps to 0.8 of a 0.5 s loop, never zero). The
one measured tolerance in the file is the shade's analytic transfer at 20%,
and the equal-power pan law needs no scenario of its own because the
multitap chapter already pinned the center at 1/√2 to 1e−12 — same code
shape, same law, cited rather than re-proven. Long-run behavior needs no
stability test at all: there is no feedback path to go wrong, which is
itself a fact the file's structure makes obvious enough not to test.
Checkpoint
A fixed bank of reels, one sacred free-running head each; record replaces
and freeze is bitwise; splices re-wrap, never rewind; the shade bypasses to
bit-transparency at the ceiling; and the composite period is the score's
arithmetic made introspectable. The promises are structural, the tests are
bitwise, and the executed raster in airport.ipynb is the human-readable
proof that the structure composes.
Events, not audio: garden.h
garden::bed recirculates events where its siblings recirculate samples,
which makes it the family's odd one out mechanically and its purest member
conceptually: the wear-as-stabilizer inversion survives the abstraction jump
intact, as arithmetic. This appendix walks the machinery — the ring, the
split between planting and firing, the chime, the quantizer, the gardener —
and the two contracts that had to be designed before they could be tested.
The event ring
Sixty-four fixed seats (std::array, nothing allocated at prepare() —
this kernel buys no tape at all), each event a pitch, a velocity, a
brightness, a position on the loop, and a plant-order sequence number. The
sequence number exists for one policy: when the garden is full, the oldest
live bloom yields to a new plant. The musical argument is stated in the
header — a touch must always speak (rejecting input makes an instrument
feel dead), and the oldest bloom has survived the most decay passes, so it
is the quietest thing on the table; retiring it is the least audible edit
available. The pinned scenario plants a distinctive high note, floods the
ring with sixty-four more, and requires the first note's pitch measurably
gone from the following pass.
Fire is not plant
note() does not sound a voice. It quantizes, seats the event at the
loop's current position, and returns; the next process() sample finds the
event's position under the playhead and fires it. The first draft did both
— plant-and-fire in note() — and the loop fired it again one sample
later, a double-trigger that fell out of the design the moment firing
became the loop's exclusive job. One mechanism, two consequences: a plant
sounds one sample late (inaudible, documented), and every sounding of every
event goes through a single code path, which is what makes the return grid
a testable promise. After each fire the event blooms: velocity times
decay, brightness times soften, retire below floor — so a bloom lives
exactly ceil(log(floor/velocity)/log(decay)) passes and the population
converges no matter the planting rate. That is the stability theorem, and
it is three lines of arithmetic instead of a saturator.
The chime
Four decaying mode doublets at the transverse-vibration ratios of the
selected material — the free-free tube's 1 : 2.756 : 5.404 : 8.933 from
the bars-and-tubular-chimes chapter of Fletcher & Rossing's The Physics of
Musical Instruments (f_n grows as (2n+1)²), or the tuned bar's
double-octave 1 : 4 : 10 : 20 from the mallet-percussion chapter — each
mode a pair of sines split a fixed few cents, the doublet splitting of a
real tube's degenerate mode pairs (same source), so the tail beats slowly
instead of decaying like a lab sine. The ratio and haste tables are indexed
[material][mode] and read at strike time, which is the whole
implementation of the material switch: instant, allocation-free, and every
live bloom re-voices at its next return. Each mode rides
its own tr808::decay_env; decay times divide by ~ratio² (radiation
damping grows with frequency), which makes the fourth mode a
tens-of-milliseconds contact tick, and scale by √(440/f) per strike, so
small high tubes ring shorter than long low ones. The upper modes scale
with per-event brightness times strike hardness (a soft strike is a dull
strike) and progressively steeply (b, b², b³), so soften strips the tick
first and mode two by exactly its ratio. Mode levels sum to at most 1, so a
chime is bounded by its velocity and the pool bound stays arithmetic; modes
above 0.45·sr stay silent rather than aliasing.
The phase rule earned a refinement when the doublets arrived: a strike on a silent tube zeroes its phases — fresh initial conditions, so the pair starts aligned and its beat blooms identically at every return, which is what keeps per-return spectral measurements deterministic — while an audible steal keeps free-running phases and glides instead of clicking. The inharmonicity moved the pitch contract rather than breaking it: the upper modes clear quickly, so the YIN oracle reads each strike in its ring-down, and the scale-contract scenario still lands every off-scale plant on the scale within 20 cents.
The tube is the identity
Two more properties hang off each pitch, and neither touches the rng. A
tube's upper modes sit up to ±3 cents off the ideal ratios — the
fundamental stays true, because a maker tunes the fundamental — and the
tube keeps a fixed seat on the stereo rack, spread scaling how far off
center. Both are drawn by tube_unit, a stateless xorshift64* hash keyed
by (fundamental-in-centihertz, index) — the metal_bank.h per-index idiom,
index 0 the seat, 1..3 the mode scatter. Stateless is the load-bearing
word: the gardener's seeded generator is never consumed, so the seed-triad
contract survives intact, the rack is identical in every instance, and
every return of a bloom rings from the same place with the same flaws.
The pinned scenarios measure the scatter by scanning a Goertzel probe
across the second mode (±0.25-cent steps resolve it), and the seat by
left/right energy share: deterministic per pitch, different across pitches,
bounded by the constants. The seat itself follows the phase rule — pan
gains snap on a silent tube and slew ~10 ms through an audible steal — and
the equal-power law is the √((1∓p)/2) form, so spread 0 makes the busses
bitwise identical (also pinned).
Quantize at entry
The scale machinery is tune.h's 12-bit pitch-class mask idiom — the
make_mask builder, the nearest-allowed search that never travels more
than a tritone — copied with citation, not included, because tune.h
reaches into tap::dsp for its detector and a garden should not link a
pitch tracker to hold five scale presets. The masks themselves are plain
public-domain scale theory, deliberately not any app's preset list.
Quantizing at entry (rather than at fire) is the semantic choice: a scale
change re-pitches nothing already planted, which keeps running gardens
stable under live tinkering and makes the contract easy to state.
The gardener and the seed
The gardener is a wind model: once the idle threshold passes, strikes
arrive on a calm/gust cycle driven by a small state machine — a gust
catches 1 to 5 neighboring tubes (gust sizes it) with 30–280 ms between
strikes, the clapper walking a few semitones per swing, and the following
calm stretches with the gust just spent so the average rate stays near one
strike per pass at any setting. Idle planting consumes the family RNG
(tr808::white_noise, xorshift64*, the seed-folding and clear-reseeds
contract) — and only idle planting does. That consumption discipline is
load-bearing: the third leg of the seeded triad, "with the gardener
disabled the seed cannot matter at all", is only true because a disabled
gardener never touches the generator. The suite pins all three legs, plus
the wind itself: at gust 1 some strikes tumble inside a gust, at gust 0
single strikes never come closer than the minimum calm (the scenario sets
decay 0 so only the gardener's own strikes are counted — planted seeds
recirculate, and returns are not wind). step_seq.h promises "no
randomness anywhere"; this kernel is the deliberate counterpoint, and the
triad is the bridge back to a reproducible test suite.
A finding: envelopes never reach zero
The return-grid scenario was first written the obvious way — the percussive
test bell surely dies between returns, so the first nonzero sample after
silence is the onset. It failed, instructively: decay_env's exponential
tail crosses the 1e−12 hard-zero more than half a second after a "20 ms"
decay, so there is no silence between returns, only −200 dB of not-quite.
The fix was to stop pretending: an instant-attack bell, an amplitude
threshold scaled to the expected return velocity, and a grid claim of
"within 8 samples" — a sixth of a millisecond — with the comment explaining
that a threshold on a sine sits a few samples into the cycle. The lesson is
general for this library: exponential envelopes make "silence" a tolerance,
and tests that assume literal zeros between notes are wrong even when they
pass.
The engineering ledger
The suite measures the output, never the internals: fundamental ratios
for the decay staircase (0.5 ± 0.05 across four returns — the fundamental,
because hardness makes whole-strike peaks fade faster than velocity, then
active_events() == 0 and the render below 1e−6), a strictly-decreasing
Goertzel sideband for softening, YIN for the scale contract, the seeded
triad rendered three times over, and structural bounds exercised at their
edges — sixty-five plants against sixty-four seats, thirty-two notes
against sixteen bells, finiteness and the k_voices amplitude bound under
sustained stealing. The two introspection counts (active_events,
active_voices) exist, as phase() does next door, so those scenarios
could be written against public surface.
Checkpoint
A fixed ring of events fired by a loop counter into a fixed pool of modal
wind chimes: plant and fire kept strictly apart, wear as per-pass arithmetic
(decay, soften, floor) with convergence as its theorem, scale masks copied
from tune.h and applied at entry, tube identity (material voicing, mode
scatter, stereo seat) as stateless hashes so nothing generative leaks into
the audio path, and a gardener whose RNG discipline makes generative
behavior compatible with a bit-exact test suite. Third
costume, same inversion: the system stays bounded because everything in it
is always fading. Every claim lives twice — garden.ipynb executed,
garden_test.cpp pinned.
Composition, not construction: tapecho.h
This is the shortest appendix in the book, and that is the point of it.
tape_loop.h was written for the Eno family — one shared header holding a
reel, a transport, and a wear stage, factored out because discreet.h and
airport.h needed the same four pieces twice. The claim implicit in
factoring it that way was that it is a library: machinery that a machine
nobody had written yet could be built out of. tapecho.h is the test of
that claim, and the result is worth recording precisely, because "we
extracted a shared header" is easy to say and rarely checked.
The result: tape_loop.h needed no changes at all. Not a new method,
not a widened clamp, not a friend declaration. A tape echo — a different
topology, a different number of read points, a different stability regime —
composed out of it exactly as shipped.
What the file actually contains
Two classes and no DSP that was not already in the library.
head is a read position with three ramps: a ratio along the tape path, a
level, and a pan. Its read() takes a reel it does not own, the motor
span, and the shared transport offset, and accumulates a panned contribution
onto the stereo busses. It is a component in the airport.h sense — a piece
the monolith is made of, reachable for testing — but honestly labeled as
not standalone-external material: a head without a reel is not a machine,
it is an index. That distinction is worth keeping straight, because the
components chapter's lesson ("the monoliths were monoliths by accident") can
be over-applied. Some seams are real and some are arithmetic.
machine owns one reel in delay-line topology, one tape::wow_flutter, one
tape::wear, and four heads. Its process() reads the heads, applies the
regeneration cap, writes the record head, and mixes. There is nothing else
in it.
The geometry: one motor
Each head's delay is span_samples * ratio - offset, where offset is the
transport error. Two decisions hide in that one line.
The first is that span is defined as the delay of a ratio-1.0 head
rather than as "the delay time", which is what makes the motor a motor: one
multiply per head and the whole layout scales together, as a tape speed
does. The alternative — per-head absolute times — would have made a speed
change into four coordinated parameter moves and lost the doppler for free.
The second is that offset is subtracted once, shared by every head. That
is physically right for a single transport (one capstan error displaces the
whole tape path) and it is also the cheap answer, so it is worth saying
plainly that the per-head phase differences of a real multi-head transport
are not modeled. It is a documented limit, not an accident.
The stability inversion, one step further
machine/tape.md derived why discreet.h may run regeneration at exactly
1.0: wear's saturator is bounded by 1/drive, so the loop is bounded no
matter the gain. That derivation does not stop at 1.0 — nothing in it does.
So this kernel lets regeneration reach k_regen_max_driven (1.5), and the
tape is bounded by |in|max + regen/drive at any setting.
The subtlety is the boundary. That guarantee exists only while the
saturator is engaged, and drive is a ramped parameter a performer can
take to zero mid-howl. At drive 0 the wear path is exactly linear with
|H| ≤ 1, so regeneration above 1.0 would grow without bound. The kernel
therefore computes the cap per sample from the current drive:
const double regen_eff = std::min(regen, (drive > 0.0) ? k_regen_max_driven
: k_regen_max_linear);
Not in the setter — in the audio path, because drive moves during
performance and a setter-time decision would be stale the moment it
mattered. The stored target keeps its high value, so pulling drive to zero
lands the loop at 1.0 and restoring drive brings the howl back. That
asymmetry between target and effective is the one piece of state in this
kernel that is not obvious from the header's public surface, which is why it
is written down twice: here, and in the file's own banner.
Why the null test is the important one
The suite's load-bearing scenario neutralizes the tape — no transport error,
no regeneration — and asserts that a one-head echo is bitwise
delay.h's Hermite multitap.
It is bitwise rather than approximate because nothing was reimplemented:
both paths compute time_ms * 0.001 * sr the same way, both clamp at the
same 2.5-sample Hermite floor, both evaluate the same polynomial at the same
fractional position, and both apply (pan + 1) * 0.25 * π to the same
k_pi. The multiply by a ratio of exactly 1.0 and the subtraction of an
offset of exactly 0.0 are both exact in IEEE-754, so the arithmetic does not
merely agree — it is the same arithmetic.
That is what makes the test meaningful. An approximate null test would pass just as happily over a second implementation that happened to be close. A bitwise one only passes if the shared code is genuinely shared, which is the proposition on trial. The notebook runs the same comparison across the C ABI so the claim also holds at the boundary the externals cross.
One consequence worth knowing when reading the test: at pan 0 the two busses
are not bit-identical to each other, because cos(π/4) and sin(π/4)
differ by one ulp in IEEE-754 doubles. That is inherited from delay.h's
pan law, it is the same in both objects, and it is exactly why the null test
compares each bus against its counterpart rather than comparing left to
right.
Checkpoint
Two classes, no new DSP, and a shared header that did not move. The motor geometry buys varispeed with one multiply per head; the regeneration cap lives in the audio path because the thing it depends on is performed; and the null test is bitwise because being bitwise is the only version of that test that proves anything.
Dice you can replay: stammer.h
Most kernels in this library are hard to get wrong quietly: a filter with a bad coefficient sounds bad. A stutter is not like that. It has three interacting integer clocks — a grid countdown, a slice origin, and a playback head — and if any one of them is off by a sample the object still sounds fine. It stutters. It grooves. It is just wrong in a way no amount of listening will surface.
So the interesting content of this appendix is not the DSP, which is a buffer and some dice. It is how you pin three clocks at once.
The pinned-dice identity
Every random draw in the kernel has a setting at which its outcome is
forced. Fire probability 1 always fires. divisions 1 always picks the
whole step. repeats 1 always plays one pass. reverse 0 never reverses.
jump 0 never reaches back. fade 0 leaves the material alone.
Set all six and the machine becomes deterministic regardless of the seed — and what it must then be is not a vague "sensible output" but a specific, checkable thing: exactly a one-step delay. At each grid point it grabs precisely the step that just went past and plays it once, so
y[i] == x[i - step + 1]
for every sample after the first grid point, bitwise.
That single assertion is worth more than three separate off-by-one tests, because it fails if the grid countdown fires a sample early, if the origin arithmetic reaches one sample too far back, if the playback head starts at the wrong index, or if any two of those are wrong in ways that would cancel in a looser test. It is also cheap to reason about, which matters: a test you cannot re-derive on a whiteboard is a test you will eventually delete instead of fixing.
The + 1 in that expression is not a fudge. The write head holds the next
write position, so after recording sample i the newest available sample
sits at position i, and a slice of length step grabbed at grid point
k·step reads positions k·step + 1 - step upward. Getting that constant
right by derivation rather than by nudging until the test passed is the
whole discipline; a test you tune to the implementation pins nothing.
Two smaller identities sit alongside it. With reverse 1 the same grab
reads end-first, so y[k·step + j] == x[k·step - j] — the mirror of the
first, which catches a reversed-index off-by-one that the forward test
cannot see. And with repeats above 1, every output sample must be either
a fresh grab or a bit-exact copy of the block one slice-length earlier;
nothing else is legal, because a slice in flight is never interrupted. That
invariant covers the repeat machinery without needing to know how many
passes the dice chose.
The draw order is part of the ABI
maybe_fire() draws in a fixed order: fire, division, repeat count,
reach-back, then the first reverse coin. Each subsequent repeat draws its
own reverse coin as it starts.
That order is not an implementation detail — it is what "a seed is a
performance" means. Reordering two draws, or adding a draw in the middle,
silently changes every render anyone has ever made with a given seed.
garden.h established the same discipline for the gardener; this file
inherits it, and the comment above the function says so in as many words so
the next person to add a parameter knows to append rather than insert.
The disabled case is the sharp end of it. At density 0 the function
returns before drawing anything:
if (m_density <= 0.0) {
return; // the dice are never rolled, so the seed provably cannot matter
}
The lazier version — draw, then compare against 0 and fail — behaves identically to the ear, and would be indistinguishable in almost any test. It would also consume one number per grid point, so the seed would matter: switch density off and on again, and the stream is somewhere else. The garden made this a family contract, and it is pinned here by a test that runs two different seeds at density 0 and requires bit-identical output.
Reading from the ring, and what it costs
A slice does not copy its material. It stores an origin and reads from the capture ring as it plays.
The alternative — memcpy the slice into a private buffer at fire time —
would be more obviously correct, and it is what a first draft wants to do.
It was rejected because it is a burst copy in the audio thread: half a
second of slice is 24,000 doubles moved inside one process() call, a spike
that does nothing for 47,999 other samples. Reading from the ring costs
nothing extra.
The price is a real failure mode, so the header states it: if a repeat train
outlives the buffered history — repeats · length + jump beyond
max_history_ms — its tail reads fresher material as the write head laps
the origin. The kernel clamps the slice length against the bought capacity
so it can never read outside the buffer, but it does not and cannot
prevent a long train from being overtaken. Sizing the history is the
caller's job, and the object argument exists for exactly that.
The envelope, and why the dip stays
Flanks are raised sine, computed to be exactly 0 at both edges and exactly 1 across the plateau, clamped per slice to half the slice so the two flanks never overlap. Repeats are sequential, not overlapped, so each junction dips to zero rather than crossfading.
Leaving it that way was a decision. An equal-power crossfade between consecutive passes is easy from here — the pieces are already in the family — and it would smooth exactly the articulation that makes a stutter read as rhythm. The dip is the transient the ear locks onto. So the file documents it as intentional, and the test asserts the exact edges, so nobody later "fixes" the dip and quietly turns the object into a tremolo.
Reuse, and the component that is not a component
The capture is a tape::reel in delay-line topology — the same class the
tape echo uses, the same class airport.h runs as a true loop. Reads here
are always at integer positions, so the family's Hermite read reduces to an
exact sample fetch; that is slightly more arithmetic than an integer index
would need, and it is kept anyway because it is one code path and because a
rate-varying sibling (tap.scrub~, planned) needs precisely this.
The randomness is tr808::white_noise, the family's seeded xorshift64*,
reached through swing_vca.h. Nothing new was written for it.
Which leaves the split: capture and slicer are separate classes under a
thin machine, per the family's components-first habit. But a slicer
needs a capture to mean anything, so — like tapecho.h's head, and
unlike airport.h's loop — it is documented as a component for
composition and testing rather than a candidate for its own external. The
components chapter's lesson is that seams often already exist and only the
monolith can reach them. The corollary, which is easier to forget, is that
not every class boundary is a seam.
Checkpoint
Three clocks, pinned by one identity: force every die and the machine must be exactly a one-step delay, bitwise. A fixed draw order, because that is what makes a seed a contract, and an early return at density 0 so a disabled generator provably cannot consume its stream. Ring reads instead of a burst copy, with the failure mode written down rather than papered over. And an envelope dip that is deliberate, tested, and therefore safe from being helpfully removed.
Two stages and a knee: fuzz.h
This appendix is mostly about two mistakes, because the DSP itself is a published recipe followed closely and there is little to explain about it that Yeh, Abel and Smith's DAFx-07 paper does not explain better. What is worth recording is what went wrong on the way, since both failures are the kind that recur.
The recipe, briefly
Two stage objects, each a conditioning highpass, a gain, a memoryless
curve, and an equalization lowpass — the paper's cascade, twice. A tone
section of three RBJ biquads outside the nonlinearity. A pedal that runs
the pair inside an oversampled region, DC-blocks, voices, and trims.
The curve is shape(x, k) = tanh(kx)/tanh(k), chosen from the family the
paper itself compares against a tabulated diode DC curve (tanh, arctan, a
tanh approximation). The normalization matters more than the choice: dividing
by tanh(k) fixes the output at full scale for unit input at every knee, so
edge changes the shape of the corner without moving the ceiling.
Mistake one: small-signal gain compounds
The tanh family's slope at the origin is k/tanh(k). It is greater than one
and it grows with the knee — 1.7 at knee 1.6, 2.1 at knee 2, 12 at knee 12.
In one stage that is a curiosity you can absorb into the gain mapping. In a cascade it multiplies. The first implementation put a fixed ×2.2 in front of a knee-3 curve, so the second stage's effective small-signal gain was about 6.6, and with the drive floor at +6 dB the limiter was already saturated with the gain knob at zero. The measured harmonic-to-fundamental ratio was 0.401 at gain 0 and 0.408 at gain 1: the knob did essentially nothing.
The reason this is worth a paragraph is that it is inaudible as a bug. The object sounded like a distortion pedal at every setting, because it was one at every setting. Only a swept measurement showed the knob was inert. The fix was to lower the drive floor below unity (−12 dB) and the second stage's fixed gain to 0.5; the ratio now runs 0.010 → 0.358.
The general lesson, stated for the next cascade someone builds here: a waveshaper's small-signal slope is part of the gain structure, and if the curve family's slope depends on a user-facing parameter, that dependence propagates to every stage downstream of it.
Mistake two: the house oversampler, and a hypothesis that was right
The oversampling chain in tap.ladder~, tap.svf~ and overdrive.h is
zero-stuff plus a 4th-order Butterworth, cut at 0.45 of the base rate
normalized to the oversampled rate. This file started as a copy of it.
Measured, that was wrong here: alias energy at the fold frequencies came out worse at 4× than at 2× (1.7e-2 against 2.8e-3). Twenty-four dB per octave leaves content just above the base Nyquist barely attenuated, and a higher factor pushes more clipper-generated content into exactly that band before decimation. Moving to 8th order improved 4× about sixfold.
It did not fix the ordering. An earlier draft of this file recorded that as an open question with one hypothesis ruled out and one surviving:
- Ruled out: numerics. At 8× the filters are cut at 0.056 normalized, where biquad poles crowd the unit circle. Tested by running the cascade's impulse response out to 400,000 samples at each factor; it decays cleanly to denormal every time.
- Surviving: imaging. Zero-stuffing by N leaves N−1 images for a single filter to suppress; residual images entering a nonlinearity intermodulate with the signal into products that are not harmonics of the input, which is precisely what the probe measures, and there are more of them at higher N.
ondes.h then supplied evidence for the survivor without being built to:
same 8th-order chain, comparably hard nonlinearity, but a source with
nothing zero-stuffed on the way up, and its sequence never reversed.
Acting on it settled it. The chain is now one 2× stage per doubling, each filtering at 0.225 of its own operating rate — a corner that never tightens however deep the cascade goes, which is the whole difference. Same probe, same material, only the resampler changed:
| tone | old 4× | old 8× | new 4× | new 8× |
|---|---|---|---|---|
| 3733 Hz | 7.4e-4 | 1.8e-3 | 2.1e-5 | 2.2e-5 |
| 4409 Hz | 7.9e-4 | 2.1e-3 | 3.7e-7 | 3.7e-7 |
| 5171 Hz | 9.2e-4 | 2.3e-3 | 2.0e-7 | 1.9e-7 |
| 6421 Hz | 5.8e-4 | 1.3e-3 | 3.9e-7 | 3.6e-7 |
| 9337 Hz | 2.9e-4 | 6.2e-5 | 1.8e-6 | 2.4e-8 |
The worst step-up past 2× is a ratio of 1.017 — flat, where the old chain ran up to 3× worse per doubling. The 2× column is unchanged in both, as it must be: one doubling is one stage either way, and that it is unchanged is the best available check that nothing else moved.
The cost is 3.16 % of a core at 8× against 3.02 % before. The filters are cheap next to the clipper they surround, which is worth knowing in advance next time this trade looks expensive.
Mistake three: one tone is not a sweep
Every number in the two sections above — the 4th-order finding, the reversal, the "2× is best" default that shipped — came from a single test tone at 3733 Hz. The tone was chosen carefully, for good reasons that are still good: it does not divide the sample rate, and its folds land where nothing else lives. It was still one tone.
Swept properly, 2× does not merely fail to be best. It collapses above about 6 kHz, and at 10499 Hz it measures worse than no oversampling at all (1.7e-1 against 1.5e-1), because the clipper's low harmonics already exceed the base Nyquist there. The default that shipped was safe only for material that stays below 6 kHz.
This is the same error as the two in the next section, one level up: those are
about choosing a bad probe, this is about choosing too few. A probe that is
correct at one point on the input domain tells you about that point. The fix
is not cleverness, it is a for loop over tones, and it costs seconds.
Two properties of this particular probe bound where the loop can go, and both are now written down next to it: it only measures folding at all above about 3 kHz, since below that harmonics 8–13 are still under Nyquist and it reads real harmonics instead; and tones that are simple rational multiples of the sample rate stack folds on top of each other or put one exactly at Nyquist, where it reads nonsense.
Two ways to measure aliasing wrong
Both were committed before being caught, and both are recorded in
fuzz_test.cpp because they are easy to repeat.
Choosing a tone that divides the sample rate. The first alias test used 3 kHz at 48 kHz. Every harmonic of 3 kHz folds back onto another harmonic of 3 kHz, so every alias hides exactly underneath legitimate content and the probes read nothing at all. The test passed happily while measuring noise. 3733 Hz puts the folds where nothing else lives.
Probing too close to the fundamental. Two of the original probe frequencies sat a few hundred Hz from a full-scale tone. What they measured was the window's spectral leakage — around 1e-3, which swamped the aliasing underneath it. Probes have to be far enough out that leakage from the loudest component is below the thing being measured.
Checkpoint
A published cascade, followed closely. One curve whose normalization keeps the knee from becoming a volume control. A gain floor set below unity because small-signal slope compounds across stages — the bug that sounded fine. An 8th-order oversampling filter because the house 4th-order one measured worse, and a cascade of 2× stages because a single zero-stuff by N was what made bigger measure worse — the imaging hypothesis, recorded as open here for two waves, then confirmed by acting on it. And a default of 4× rather than 2×, because the 2× default had been generalized from one test tone and collapses above 6 kHz.
One tape, two read patterns: scrub.h
When stammer.h shipped, its header made a promise in its own limits
section: slices play at ±1 rate, a performable pitch-bending playhead over
live capture is a different object, and sharing this capture is the plan.
This file is that promise being kept, and it is worth recording that the
sharing turned out to be literal. scrub.h includes stammer.h and uses
stammer::capture itself — one tape_loop.h reel under an advancing write
head — rather than keeping a second copy of the same idea.
The only thing the stutter had to grow was capture::read_frac, a
fractional Hermite read. Its ±1-rate slices never needed one.
The rest of the file is two classes: head, the grain scheduler, which owns
the grain pool, the hop clock and the spray dice and reads a capture it does
not own; and machine, which is one capture, one head, the freeze gate, the
drift and the balance. Same parts-then-composition habit as tapecho.h and
stammer.h, for the same reason: head is a read pattern, not a machine,
so it is testable and composable without being an external.
The defect that measurement caught and nothing else could
The first cut anchored every grain at the position. That is the obvious thing to do — the position is where the user is pointing — and it is wrong in a way that is genuinely hard to hear.
Here is the mechanism. If every grain's origin is write_head − lag, then
origins advance at the write head's speed, which is exactly 1. Each grain
then plays from its origin at rate. Inside a grain the pitch is correct.
Across grains the average read rate comes back to 1, because the origins
reset it every hop.
So a steady tone comes out at its original pitch, with a comb of grain-rate sidebands around it. The pitch knob did not transpose. It added texture, and texture is what you expect from a granulator, which is exactly why no amount of listening was going to find this.
The fix is a phase-continuous read head: the origin advances at rate, and
is pulled back toward the position only once it has wandered more than
±1.5 grains. Every pull-back is a splice, which is the cost, and the bound
is chosen by sweep rather than taste. Band energy retained around the
transposed pitch, at wanders of ±0.5 / ±1 / ±2 / ±3 / ±4 grains:
| wander (grains) | mean | worst |
|---|---|---|
| ±0.5 | 0.933 | 0.716 |
| ±1 | 0.958 | 0.820 |
| ±2 | 0.965 | 0.874 |
| ±3 | 0.990 | 0.918 |
| ±4 | 0.993 | 0.940 |
Flat past 3, and every extra grain of wander is a grain of position error,
so k_wander_grains = 3.0.
The unity case is special-cased to zero error rather than accumulated,
which is what keeps the null exact: at rate == 1 there is nothing to
wander from.
Measure the band, not the bin
This is the second thing worth carrying out of this file, and it nearly inverted the conclusion above.
A single-bin probe reads the fixed kernel as badly broken. The splices spread the transposed partial into a comb a few hertz wide; a rectangular-window Goertzel sitting on one line saw 0.02 where the band figure was 0.43. Had that been the first measurement taken, the fix would have looked like the bug.
Measured properly — energy in a ±15 Hz band around the transposed pitch, against the same band of a perfect shifter — 98.8 % lands where it should, worst case 91.7 %. What the splices cost is concentration, not pitch: 92.0 % as focused as a clean shift, 75.0 % at worst.
The general rule, stated for the next time someone here measures a pitch-shifter: if the process can smear a partial, a single-bin probe is measuring the smear, not the partial. Integrate a band wide enough to contain the artifact you already know about.
And then, immediately, the same mistake in its other half. The comparison
against tap.pitchaccum~ used that ±15 Hz band unchanged across the whole
sweep — but ±15 Hz is about 115 cents wide at 220 Hz and only 26 cents at
932 Hz, so at the top of the sweep the probe was again narrower than the
process it was measuring, and it produced two readings of 0.0001 and 0.0006
that were recorded as near-total cancellations of a shipped object. Widened
to a constant 3 %, they read 0.63 and 0.85 and no cancellation exists. The
retraction and what survives it are issue #33.
So the rule has a second half: a band wide enough in the units the process works in. A pitch shifter works in cents. A fixed hertz window is a different width at every pitch, and the place it is narrowest is exactly where a shifter's error is largest.
Two related mistakes are recorded here because both were committed:
- Analysing mostly silence. The first wander sweep ran 1 second of material with a 900 ms position lag, so most of the analysed window was tape that had not been written yet. Extended to 3 seconds, analysing the last third.
- Feeding a discontinuity into the test. A slew test drove the object
with a sine and then, mid-test, called
process(0.5)with a literal DC sample to change a parameter. That step was an input transient, and the 0.48 jump it produced was the test's own fault. Continuous tone index, and the same bug was then fixed pre-emptively indiffuseur_test.cpp.
The null, and the arithmetic that makes it exact
Hann satisfies constant-overlap-add at hop = size/overlap, so the window sum
is exactly 1 at overlap 2 and above, and normalization is 2/overlap so the
level holds across settings. With pitch at unity, spray at zero and the
position on a whole sample, the object is the input delayed to 4.4e-16.
It is exact only when size divides evenly by overlap, because the hop is
an integer number of samples; otherwise a small periodic ripple survives in
the window sum. It is inaudible at musical sizes, and it is why the null
test chooses the numbers it does (480 samples of lag, 96 of size) rather
than round milliseconds.
The mix control needed the same care as the diffuseurs' — an equal-power
blend written as cos/sin does not return exactly zero at the endpoint,
and a wiring null that reads 6.1e-17 instead of 0 is not a null. Both ends
are short-circuited exactly.
The grain pool starves rather than steals
Shrinking size sharply while grains are in flight can leave every slot
busy at the moment the next grain is due. That grain is dropped, not
allocated by stealing a slot from a grain mid-window, because a steal cuts a
Hann window in half and clicks. The audible cost is a momentary dip, bounded
by the pool being two slots deeper than the maximum overlap.
A limit that is not fixed, on purpose
A grain born lag samples behind the write head and playing at rate r
reaches lag − size·(r−1) behind it by its end. Transpose up with the
position near the live edge and the grain's tail runs off the front of the
tape into the oldest material.
Nothing clamps this. Clamping would silently bend the pitch to keep the
grain in bounds, which is a worse failure than the seam — the object would
stop playing the interval you asked for and never say so. The constraint is
documented (keep the position at least size·(rate−1) back) and left to the
player.
Checkpoint
One capture, shared literally with the stutter, plus one fractional read that the stutter did not need. A phase-continuous read head, because anchoring grains at the position quietly cancels the transposition — the defect of this file, invisible to listening and obvious to a sweep. A wander bound measured rather than chosen. And a measurement lesson worth more than the kernel: a single-bin probe on a smeared partial reads the fix as the bug.
Driven, not struck: diffuseur.h
This file exists because a plan was wrong in a useful way. The Ondes family
plan said the diffuseurs would inherit garden.h's modal machinery, and
they do — mode ratios, doublet splitting, per-mode decay. What it did not
say, and what reshaped the file, is that a diffuseur is driven. There is
no trigger here and no decay_env. The input excites the body continuously
and the body rings at its own rates, which is grm_comb.h's situation
rather than the chime's.
Five classes: mode, plate, sympathetic, harp, transducer, and two
cabinets over a shared cabinet base. Nothing else.
Unit peak gain, and everything it saves
mode is the constant-peak-gain two-pole resonator (Steiglitz; Smith,
Introduction to Digital Filters): poles at radius R, zeros at ±1, and
b0 = (1 − R²)/2.
That choice pays three times, and it is worth spelling out because it is the reason this file has almost no defensive code in it.
- Peak gain is 1 at any Q. So a bank of weighted modes is bounded by the sum of its weights. The plate's eight weights sum to exactly 1, which means the body cannot output more than its input, and there is no limiter anywhere in the file.
- Changing
decaydoes not change the level. With a plain two-pole resonator, moving R moves the peak gain, so a decay knob is also a volume knob. Here it is not. - The zeros at ±1 are exact nulls at DC and Nyquist. So there is no DC blocker on the body either. It cannot accumulate one.
None of that is novel — it is a textbook resonator used for the reason the textbook gives — but the cumulative effect on a file that runs sixteen of them plus twelve delay loops is large.
The order is the argument, and it is a bitwise test
The instrument's signal reaches the transducer first, and the transducer's motion excites the body. So the nonlinearity is upstream of the resonator.
That is the central design claim of the file, so it is pinned rather than
described: a scenario builds transducer → plate by hand and checks that a
whole metallique is bitwise identical to it, and that the reverse
wiring — resonate, then distort — differs by 28 % of peak.
Getting that null to be actually bitwise took one fix. cabinet::blend is
an equal-power crossfade written with cos/sin, and cos(π/2) in double
precision is 6.1e-17, not 0. A wiring test that reads 6.1e-17 has not
demonstrated identity; it has demonstrated approximate identity, which is
the thing the test exists to distinguish from. Both ends of the blend are
now exact short-circuits: mix 0 returns the dry input bit for bit, mix 100
returns the wet.
The transducer's bound is 2/saturation, not 1/saturation
The moving-iron model squares the drive, and vca::swing_shape bounds the
result at 1/saturation. The obvious test — output stays under
1/saturation — failed at 1.49 against a bound of 1.25.
The test was wrong, not the code. A hard-driven squared law produces a nearly-constant positive waveform: it sits up near the ceiling and dips toward zero. Removing its DC recentres that, so the excursion below the mean adds to the excursion above it, and the worst-case swing after the DC blocker is up to twice the saturator's own bound.
The corrected bound is 2/saturation, documented in the header, and the
scenario now asserts both sides of it — greater than 1/sat, less than
2/sat — so the test still catches the saturator disappearing entirely.
The general shape of this mistake is common enough to name: a DC blocker after an asymmetric nonlinearity is not free. It does not just remove an offset; it converts an offset into headroom you have to have.
A measurement that measured its own edges
The palme's selectivity scenario drives the board with a tone and measures what is still ringing after the tone stops. First version: switch the tone on, switch it off, measure the tail. It failed — 3.66× selectivity against the 4× asserted — and the failure was real but not about the strings.
Switching a tone on and off is a step, and a step is broadband. It excites every string on the board, so the tail contained twelve strings ringing regardless of what frequency had been played. The measurement was reading its own edges.
Fading the drive in and out over 250 ms removes the step. Every one of the twelve strings then passes, with the worst at 4.4×. The same fade is what the book figure uses, and the figure's caption says so, because a reader reproducing it without the fade will get the wrong answer.
Twelve strings
Widely copied hobbyist build pages describe the palme as two banks of twelve
strings. The peer-reviewed source (Wijnand, Boutin, Jossic & Maniguet, Forum
Acusticum 2023) says twelve, and k_strings = 12 with a comment saying
which source won and why.
Their tuning is not published anywhere found, so it is a parameter rather than a constant, and the header says that too. Guessing a tuning and hard-coding it would have been the same category of error as the twenty-four.
Where recreation begins
The instruments, their dates, their excitation and their transducer type are peer-reviewed. The modal data is not — no ondes-specific measurement of either body exists in any of the four sources read — so the plate uses Fletcher & Rossing's free circular plate (Rayleigh's Chladni ratios at Poisson 0.3) and the strings use the harmonic series.
This is stated in the header at the top rather than in a limits section at
the bottom, because it changes what the object is: a recreation of the
general physics, not a model of Martenot's instruments. The same applies to
asymmetry and saturation — the source establishes that the moving-iron
driver is nonlinear and that Thiele–Small does not describe it, and then
does not hand over a curve. Those two coefficients are voiced by ear and
labelled as voiced by ear.
A diffuseur with both at 0 is a linear resonator and is missing a real stage. That is a choice the caller may make, and the header says so rather than forcing a minimum.
Checkpoint
A textbook resonator chosen for three properties that between them remove
the limiter, the DC blocker and the decay/level coupling. A bitwise null
that pins the transducer upstream of the body, which required making an
equal-power blend exact at its endpoints. A saturator bound corrected from
1/sat to 2/sat because a DC blocker after an asymmetric nonlinearity
buys headroom, not just centring. A selectivity test that had to stop
measuring its own on/off step. And a provenance line drawn where the
published sources actually stop.
A citation, an identity, and a sign: ondes.h
Three classes — triode, detector, voice — and three things worth
recording about how they got here. One stage turned out to need no design
decisions at all. One approximation turned out to be an exact identity. And
one sign error made a distortion knob run backwards.
The circuit is Najnudel, Hélie, Roze & Boutin, "Simulation of an ondes Martenot circuit", IEEE/ACM TASLP 28, 2651–2660, 2020, modelling instrument No. 169 as five port-Hamiltonian stages. This file is not that: their full solve runs at 768 kHz and their plugin costs 85 % of a laptop core. What it takes from them is their own published reductions plus their published component values, and the header says which is which.
The tube is a citation, not a design
The plan framed the valve stage as a choice: a published grid-conduction curve, or the tanh family with an asymmetry bias voiced by ear. It is neither, and finding that out took nothing more than reading the paper properly.
The paper names a tube model — the enhanced Norman Koren model (Koren, Glass Audio 8(5), 1996, with Cohen & Hélie's grid-current branch, AES 129, 2010) — writes out its three equations, and publishes parameter sets in Table II fitted to the actual valves in ondes No. 169, together with each stage's supply voltage, cathode resistor and plate load.
So there was nothing to voice. k_6f5, k_6c5, k_2a3, k_op_demod,
k_op_preamp and k_op_power are Table II transcribed, and the header says
they are the citation.
A stage is then the static solution of ipc(vpc, vgc) = (Vbias − Vk − vpc)/Rp on the load line, with cathode bias Vk = Rk·Ipc found at the
quiescent point. That is a memoryless nonlinearity in exactly the
DAFx-07 sense, which matters for a practical reason: tabulating it is not an
approximation of the model, it is the model. The table is rebuilt on a
tube or operating-point change and read with linear interpolation, so the
audio path costs a lookup rather than a root find.
The published points bias sanely — the 6C5 demodulator lands at Vk 2.70 V, Vp 86.5 V, Ip 2.70 mA, gain 4.86 — which is its own small confirmation that the transcription is right.
The sign that made the drive knob run backwards
The stage must invert, as a real common-cathode stage does, and this is load-bearing rather than cosmetic. The valve's asymmetry acts on whichever side of the waveform reaches its grid. An early cut normalized the output by the signed small-signal gain, which quietly un-inverted the stage, so the curve's lopsidedness landed on the wrong half of the waveform.
The symptom was unambiguous once measured: turning drive up reduced
total harmonic content. A distortion control that gets cleaner as you push
it is not a subtle bug, but it is only visible in a sweep — at any single
setting the object sounded like a valve.
Two changes fixed it. The curve is now the true (inverting) plate swing, and
normalization is by the gain's magnitude. And voice::core applies the
demodulator's own grid-leak inversion explicitly — a growing envelope drives
that grid toward cutoff — so the two inversions put the demodulator's plate
in phase with the envelope while the curve has meanwhile acted on the
underside. drive now sweeps harmonic content 0.221 → 0.344, monotonically.
The gain-staging lesson from fuzz.h was applied here from the start rather
than learned again: each stage is normalized by its own small-signal gain,
so drive changes the distortion and not the level.
The detector is an identity, not a simplification
The plan's instruction for this stage was "synthesize the difference tone directly as a sinusoid", and catching that as a mistake is the most valuable thing this build did.
The paper's 0.03 % distortion figure and its licence to replace oscillators
with a sinewave generator apply to the oscillators. The demodulator is
not a mixer handing you a difference tone; it is an envelope detector, and
the envelope of cos(Φ) + cos(Φ − φ) is 2|cos(φ/2)|, whose Fourier series
puts H2 at −14.0 dB, H3 at −21.3 dB and H4 at −26.4 dB. Synthesizing a
sinusoid would have discarded the instrument's largest single source of
harmonics before any of the modelled stages ran.
What replaces the carrier is better than a simplification. For amplitudes 1
and depth, the envelope is exactly
sqrt(1 + depth² + 2·depth·cos(2π f t))
so the 80 kHz carrier drops out of the arithmetic rather than being
approximated away. Running the published RC detector on that closed form —
instant attack through the diode, 200 µs decay through R4·C21 — reproduces a
full heterodyne-plus-diode-plus-RC simulation to within 0.10 dB on every
harmonic at every pitch tried (ondes.ipynb §2).
There is one systematic difference, and it is worth knowing it is systematic rather than noise: the closed form sits a uniform 3.0–3.2 % high, because a follower chasing real carrier half-cycles never quite reaches the peak between them. On a synthesizer with a level control, that is a constant.
The detector's characteristic pitch dependence comes along free, out of the same 200 µs: H2 runs −14.0 dB at A2 to −19.3 dB at A6, and the level falls 2.0 dB across those five octaves.
And a bonus nobody planned: because the closed form is parameterized by the
two oscillator amplitudes, oscillator balance becomes a physical timbre
control. depth is a real mismatch between two real oscillators, not an
invented knob.
Three measurements that lied, and what they were doing
All three were committed to a notebook or a header before being caught.
Too few periods. The first measurement of the detector's harmonics at low pitch used a window holding about 2.75 periods of the fundamental. Spectral leakage at that resolution dominated everything, and it produced a confident, wrong claim in the header: "−9.8 dB at A2, level falls 9.7 dB". Redone with 60 cycles, the real answer is −14.0 dB and 2.0 dB. Both numbers were in a shipped header before the recheck.
Probing where the answer is exactly zero. The aliasing scenario probed
half-integer harmonics of a tone that was exactly periodic in the analysis
window. Those bins are analytically zero, so it measured −281 dB and passed
triumphantly. Fixed by computing the actual fold frequencies for a tone at
2637 Hz — deliberately not a submultiple of 48 kHz — and skipping folds that
land near real harmonics. This is the same family of error fuzz.h records
under "choosing a tone that divides the sample rate", committed again in a
different disguise.
Stopping the sweep at the first plateau. The header initially claimed "4× is the knee, then flat". The notebook's own more careful run — settled state, 131072-point Hann — showed 8× continuing to improve in the top octave. Corrected to "never worse", with the full table in the header, the test comment and the notebook.
The evidence that closed an open question in fuzz.h
fuzz.h measured its oversampling sequence going the wrong way — 4× worse
than 2× — and had left an untested hypothesis behind: that the culprit is
imaging, since zero-stuffing by N leaves N−1 images for one filter to
suppress, and residual images entering a nonlinearity intermodulate into
products that are not harmonics of the input.
This file runs the same 8th-order Butterworth chain around a comparably hard nonlinearity, and its sequence never reverses:
| tone | 1× | 2× | 4× | 8× |
|---|---|---|---|---|
| 587 Hz | −79.3 | −91.2 | −104.5 | −103.8 |
| 1175 Hz | −65.8 | −77.2 | −90.6 | −92.5 |
| 1760 Hz | −57.6 | −70.9 | −81.1 | −82.2 |
| 2637 Hz | −51.1 | −61.4 | −71.8 | −83.8 |
| 3520 Hz | −45.4 | −56.8 | −67.0 | −74.2 |
The difference between the two files is exactly the hypothesis: this object is a source. Nothing is zero-stuffed on the way up — the detector simply runs fast — so there are no images at all.
That was evidence, not proof — the nonlinearities differ too, and one confounded comparison does not settle a question. But it was the first evidence either way, and it pointed somewhere specific enough to act on.
Acting on it settled it. fuzz.h now cascades one 2× stage per doubling
instead of zero-stuffing by N once, each stage filtering at a corner that
never tightens however deep the cascade goes. Its reversal is gone — worst
step-up past 2× is a ratio of 1.017 — and its 4× and 8× improved by two to
four orders of magnitude, for about 5 % more CPU. This file needed no change,
having no upsampler to fix.
Worth naming the shape of it, because it is not the usual one: the evidence that resolved a two-wave-old open question in one file came from building a different file that happened to differ in exactly the right variable. It was not designed as an experiment. It was noticed, written down in both headers as evidence rather than proof, and left where the next person would trip over it.
A wrapper test that found a kernel bug
tap.ondes~'s Min-level test asserts something a patcher would otherwise
file as a bug report: with the key at rest, the object is exactly silent.
It failed.
voice::set_smooth_ms set the voice's own ramps but never forwarded to
touche::key, which keeps its own slew. So a key sitting at zero with
@smooth 0 still sounded for 20 ms after every parameter touch.
This is the two-layer split working the way it is supposed to. The kernel suite tests DSP promises; the wrapper suite tests what a patcher will actually observe, and those are not the same set. The fix landed in the kernel with its own scenario, not in the wrapper.
What this file will not do
The real instrument has switchable waveform registers. Their filter shapes are in none of the sources obtained. Adding them from imagination is the one thing this file is careful not to do, and the omission is stated in the header, the object description and the reference page rather than left as a gap someone might charitably fill later.
Two controls are choices — where the intensity key sits (keyplacement)
and the coupling transformer's winding sense (polarity) — because the
paper's five stages do not settle either. Both are labelled as choices, and
both were measured to confirm they are audible ones: about 0.09 and 0.12 of
total harmonic content respectively.
Checkpoint
A stage that required no design because the paper published the model and
its fitted parameters. A detector that is exact rather than approximate, and
cheaper than the thing it replaces. One sign error that inverted the meaning
of a distortion knob and was invisible at any single setting. Three
measurements that lied in three different ways, all recorded. The evidence that closed fuzz.h's
oversampler question, from a file that happened to differ in exactly the
right variable and was not built as an experiment. And a wrapper test that found a kernel bug,
which is the split doing its job.
How to read a recipe
The first parts of this book keep two promises: the object chapters say what each tool is for, and the machine chapters say why to trust it. This part makes a third kind of promise. A recipe puts several objects on one patch cord and chases a specific sound — a record you have heard, an instrument you have coveted — and tells you honestly how close the kit gets.
Recipes are held to the house rules, adapted:
- Every knob named exists, spelled the way the attribute is spelled. A
recipe is checkable against the reference pages; if it says
@decay 0.8, that attribute takes that value on the shipping object. - Settings are starting points, not measurements. A recipe's numbers get you into the neighborhood; your ears walk the last block. Where a chapter number is a measurement (a decay time, an alias floor), it still cites the executed notebook or pinned test that carries it — the recipes borrow those numbers rather than re-deriving them.
- Provenance stays honest. When a recipe chases a record, it says what is documented about how that record was made and what is folklore. When it chases an instrument, it leans on the same published analyses the kernels were built from. What a recipe never does is claim to be the record — mix, room, tape, and hands are not in the box.
- Every recipe ranks its ingredients. The house habit from the Moog recipe in the oscillator chapter: list what each element buys, in order of importance, so you know what to cut first when CPU or taste says so.
Each recipe has the same skeleton: the sound and where it came from, the signal chain, the settings (tables for knobs, grids for patterns), what each ingredient buys, and — because every tool is sometimes the wrong tool — when to leave the recipe and cook something of your own.
One machine, four decades
The TR-808 sold poorly, was discontinued in 1983, and then spent forty years becoming the most influential drum machine ever built — not by being realistic, but by being itself in four different genres' hands. This recipe visits four of those hands: the 1982 electro of "Planet Rock," the same year's slow soul of "Sexual Healing," the tuned-kick boom of Miami bass, and the half-time rolls of trap. Same eight circuits every time; what changes is the pattern, the accents, and which knob someone dared to turn all the way up.
One honesty note before the first grid: these patterns are starting
points, not transcriptions. Where a record's production story is
documented, the recipe says so; the grids themselves are the versions ears
agree get you into the neighborhood, and your ears finish the trip. The
voice knobs, on the other hand, are exact — every attribute below is
spelled as the shipping object spells it, and the calibration numbers
behind the voices live in the drum machine chapter and the
tr808_calibration.ipynb
notebook.
The scaffold every recipe shares
One phasor~ is the transport; every tap.808.seq~ row reads it; every
row's output cable is a voice's trigger input. The phasor's frequency for a
16-step bar of 4/4 is BPM ÷ 240 (four beats per cycle, four sixteenths
per beat). Rows fed the same ramp are sample-locked forever — that is the
sequencer's phase-derived design (see its machine
chapter), and it is why nothing below mentions sync.
phasor~ (BPM/240)
├── tap.808.seq~ ──▶ tap.808.kick~ ──┐
├── tap.808.seq~ ──▶ tap.808.snare~ ──┤
├── tap.808.seq~ ──▶ tap.808.hat~ ──┼──▶ +~ ──▶ tap.limi~
├── tap.808.seq~ ──▶ (open) hat inlet 2┤
└── tap.808.seq~ ──▶ tap.808.cowbell~──┘
Program a row with two lists: hits (which of the 16 steps sound, 1/0 per
step) and accents (which sounding steps lean, 1/0 per step). An accented
step emits the row's accented level (default 0.5), a plain step emits
plain (default 0.01) — those defaults are the hardware's accent knob at
noon, and they matter more than they look, because a voice's trigger
amplitude is a voltage on the 4–14 V bus: an accented hit is punchier and
differently voiced, not merely louder. Raise accented toward 1.0 when a
groove should hit like the accent knob cranked. Single steps tweak with
step <n> <velocity> (1-based), and each row's 16 slots (store/recall)
hold your fills.
In the grids below, X is an accented hit, x a plain one, . a rest.
1982, the Bronx via Düsseldorf: the "Planet Rock" kit
The documented part: Afrika Bambaataa and producer Arthur Baker built "Planet Rock" on a rented TR-808, borrowing Kraftwerk's melodies, and its kit — dry kick, clap-snare backbeat, offbeat cowbell — became the electro sound. The orchestra stabs were a sampler's; everything percussive is the machine's.
step: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
kick: X . . . . . . x . . x . . . . .
snare: . . . . X . . . . . . . X . . .
clap: . . . . X . . . . . . . X . . .
closed: x . x . x . x . x . x . x . x .
open: . . . . . . . . . . . . . . x .
cowbell: . . x . . . x . . . x . . . x .
- Tempo ≈ 129 BPM →
phasor~ 0.5375. tap.808.kick~:@decay 0.35 @tone 0.55— the electro kick is short and clicky, not the boom (that comes later in this chapter).tap.808.snare~:@tone 0.6 @snappy 0.7; layertap.808.clap~on the same backbeat row — the clap-plus-snare composite is half the sound.tap.808.cowbell~on the offbeats,@level 0.6. The drum machine chapter's line stands: more cowbell is a patching decision.- Hats: closed 8ths; the open hat answering just before the bar turns.
- Fill:
store 1the main pattern, program the classic descending-tom fill (tap.808.tom~,@size high→mid→lowon three rows) into slot 2, andrecall 2a bar before the phrase ends —quantize cycle(the default) swaps it exactly on the downbeat.
1982, Ostend: the "Sexual Healing" slow jam
The documented part: Marvin Gaye programmed the TR-808 himself for "Sexual Healing," and it became one of the first major hits carried by the machine — proof in the same year as "Planet Rock" that the same circuits could whisper. The kit is soft, sparse, and riding the plain/accent distinction rather than density.
step: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
kick: X . . . . . . x . . . . . . . .
snare: . . . . x . . . . . . . x . . .
closed: x . x . x . x . x . x . x . x .
open: . . . . . . x . . . . . . . x .
claves: . . x . . . . . . . x . . . . .
- Tempo ≈ 94 BPM →
phasor~ 0.3917. tap.808.rim~ @model claves— the high tick is the hook of the kit.@level 0.5keeps it a seasoning.tap.808.kick~:@decay 0.6 @tone 0.35— rounder than electro, still polite.tap.808.snare~:@snappy 0.35 @tone 0.4— more drum, less noise.- Leave the sequencer's
plainlevel at its 0.01 default and place accents sparingly; at this tempo the difference between a 4 V hit and a half-accented one is the entire feel. - A touch of
@swing 0.15on the hat row loosens the grid the way a human thumb on the start button did.
Late eighties, Miami: the kick is the bassline
Miami bass turned the kick's decay knob to the top and discovered the
808's bass drum is a tuned instrument — a bridged-T resonator whose
fundamental sits near 49 Hz (measured within 2.4 % of a real unit across
the knob grid; see the calibration pass in the drum machine chapter). Turn
decay up and it rings for seconds; give two copies two tuning ratios
and you have a two-note bassline.
step: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
kick A (root): X . . . . . . . . . X . . . . .
kick B (fourth): . . . . . . X . . . . . . X . .
snare: . . . . X . . . . . . . X . . .
closed: x . x x x . x x x . x x x . x x
- Tempo ≈ 126 BPM →
phasor~ 0.525. - Two
tap.808.kick~objects, two rows.tuningis a ratio of the stock fundamental, so target Hz ÷ 49 ≈ your setting: kick A@tuning 1.0(G1, where stock already sits), kick B@tuning 1.33(≈ C2, the fourth). Both@decay 1.0— the whole genre is that knob at the top. @tone 0.2keeps the click out of the way of the ring;@attack 0.5softens the punch mechanism if the notes should bloom instead of hit.tap.808.snare~ @snappy 0.8 @drive 6— the swing-VCA drive is the crack that cuts through the sub.- Watch the sum: two ringing kicks stack.
tap.limi~on the bus is the modern answer; ridinglevelper voice is the period one.
The 2010s: trap, and the arithmetic of rolls
Trap keeps Miami's tuned, sustained kick and moves the snare to beat 3 —
the half-time frame — then spends all its rhythmic budget on hi-hat
subdivision games. Those games are where this sequencer's phase-derived
design pays off: rows of different lengths off one phasor divide the same
bar differently, so a 32nd-note roll row and a 16th-triplet roll row are
just length 32 and length 24 — polymeter as arithmetic, measured in the
sequencer notebook.
step: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
kick: X . . . . . . x . . x . . . . .
snare: . . . . . . . . X . . . . . . .
closed (len 16): x . x . x . x . x . x . . . . .
roll (len 32): steps 25–32 hit, plain (32nds on beat 4)
triplet (len 24): steps 19–24 hit, plain (16th triplets, beats 3–4)
- Tempo ≈ 140 BPM →
phasor~ 0.5833. - Tune the kick to the song's key with the ratio table: E1 ≈
@tuning 0.83, F1 ≈0.88, G1 =1.0, A1 ≈1.12.@decay 1.0 @tone 0.15, and keepsighat its default 1.0 — the pitch relaxation is the 808-bass glide everyone samples. - The roll rows: program hits only on their last steps (as above), leave
them muted (
@mute 1), and unmute for the bar that needs the roll — or keep separate patterns in slots andrecall. Fast rolls do not machine-gun: the voices' filter states persist across triggers, so a roll interferes with the ringing tail like hardware (pinned by the family's tests; see the drum machine chapter). - Alternate hat voicing per unit:
@seedis which 808 you own, and@tolerance 0.3puts the metal bank's oscillators off-grid the way resistor variance really does (Werner et al. measured up to ~20 % — the chapter has the numbers). Two hat objects, two seeds, panned, is a stereo kit for free.
What each ingredient buys
- The pattern and its accents. Four decades of genre difference above is mostly the grids. The accent flags are not dynamics polish — they are the hardware's second voicing per drum. Spend your time here.
decayandtuningon the kick. One knob separates electro from Miami; one ratio puts the kick in the song's key.- The composite backbeat. Clap + snare on one row (electro), or snare
drive(Miami, trap) — the backbeat carries the genre signature after the kick. - Polymeter rows for rolls. Two extra rows, two
lengthvalues, and the trap chapter of the machine's biography writes itself. seed/toleranceon the metal. Seasoning, in the salt sense: invisible until you A/B two units.
When to leave the recipe
- You want those records, exactly. Mix, tape, room, and a human on the start button are not in the box; at some point the honest tool is the actual sample.
- You want velocity-per-step expression. The
velocitieslist gives a row continuous 0..1 levels — but note it trades away the two-level hardware model; the accent bus is the 808's own idiom. - You want 909, LinnDrum, or DMX. Different circuits, different machines — this family models one instrument, and its refusal to be generic is the point.
Three oscillators into a ladder
The oscillator chapter ends with the Moog recipe's core — the three-voice saw stack and the driven ladder — and ranks what each ingredient buys. This recipe finishes the instrument: the two envelopes, the amplifier, the gate, and the settings that turn one signal chain into the two patches everyone actually means by "Moog" — the bass that walks and the lead that sings. The model here is the classic three-oscillator monosynth voice: three oscillators into a mixer, one four-pole ladder, one loudness contour, one filter contour, glide on the pitch. Nothing below requires an object the package doesn't ship. And once the voice stands, the next recipe drives it at the records with names on them — Winwood, Worrell, Wright, Emerson.
Companion material: the oscillator chapter (the stack's
rationale and the analog section's ranges), the ladder
chapter (every filter number below is measured in its
notebook),
and the reference pages for tap.adsr~ and tap.vca~.
The voice, wired
pitch (midi note) ──▶ mtof ─┬─▶ tap.vco~ (voice 1)──┐
├─▶ tap.vco~ (voice 2)──┼─▶ *~ 0.36 ─▶ tap.ladder~ ─▶ tap.vca~ ─▶ out
└─▶ ÷2 ─▶ tap.vco~ (3)──┘ ▲ ▲
gate (0/1 signal) ──┬───────────▶ tap.adsr~ (filter contour) ─▶ *~ amount ─▶ +~ base ──┘│
└───────────▶ tap.adsr~ (loudness contour) ─────────────────────────┘
- Pitch arrives as note frequencies (floats into each
tap.vco~left inlet; halve for voice 3's octave-down). The oscillators' ownsmoothramp is the glide knob — no portamento object exists or is needed. - Gate is any signal that rises above
tap.adsr~'sthresholdand back — the envelope reads the gate by level, per sample. Atap.303.seq~gate output (1.0 plain, 2.0 accented) drives it directly, which also gets you slides for free; so does a MIDI-driven 0/1 signal, or thetrigger 1/trigger 0attribute messages for mouse-driven patching. The defaultmode analoggives the envelopes below the RC curves a Model D actually had;velocity(off by default) lets the gate's amplitude scale the hit. - The filter contour scales into the cutoff's signal inlet: envelope ×
amount(Hz) +base(Hz) intotap.ladder~'s right inlet. The classic panel's "amount" knob is your*~. - The loudness contour multiplies the ladder's output —
tap.vca~with the envelope into its gain inlet keeps the option of@circuit warmsaturation later.
One period-correct honesty note: the original panel's contours are
attack/decay/sustain with a release switch (release equals decay, on or
off). tap.adsr~ gives the full four stages; set release equal to
decay and you have the switch's "on" position.
The stack and the ladder
The three-voice table is the oscillator chapter's, reproduced so this page patches alone:
| voice | frequency | detune | drift | seed |
|---|---|---|---|---|
| 1 | f | −4 | 8 | 11 |
| 2 | f | +5 | 8 | 22 |
| 3 | f ÷ 2 | +2 | 10 | 33 |
All three: @shape 2 (saw), @jitter 3 @track 2 @imperfect 0.3, and
smooth per the patch below. Sum through *~ 0.36 (≈ 1/2.8, headroom for
three voices), then tap.ladder~ at the chapter's voicing:
@mode lp24 @resonance 0.35 @drive 9 @asym 0.45 @comp 0.25. Keeping comp low
preserves the authentic passband droop; drive 9 sits where the ladder
notebook measures the tanh stages just starting to thicken (3.5 % THD at
8 dB). Spend the character budget in the filter first — the chapter's
measurements are the argument.
Patch one: the bass
The left hand of a decade of records: short filter contour, no vibrato, glide short enough to read as punch rather than portamento.
| control | setting |
|---|---|
all tap.vco~ smooth | 25 ms |
filter tap.adsr~ | @attack 2 @decay 220 @sustain -18 @release 220 |
filter amount / base | 2500 Hz / 120 Hz |
ladder resonance | 0.25 |
loudness tap.adsr~ | @attack 2 @decay 400 @sustain -3 @release 120 |
The sound lives in the filter contour's decay: 220 ms is the "wah" that
articulates each note. Shorten toward 120 ms and it turns percussive;
lengthen toward 400 ms and it turns brassy. For a rounder, more
sub-friendly bass, drop drive to 3 and asym to 0.2 — the even
harmonics are lovely on a lead and muddy on a bass amp. If anything
downstream cares about DC, remember the ladder chapter's warning:
an asymmetric saturator can leave a small signal-dependent offset —
tap.dcblock~ after the VCA is one object of insurance.
Patch two: the lead
The singing version: longer glide, opened filter, resonance high enough to color but under the edge, and the release switch "on."
| control | setting |
|---|---|
all tap.vco~ smooth | 80 ms |
filter tap.adsr~ | @attack 15 @decay 600 @sustain -8 @release 600 |
filter amount / base | 4000 Hz / 300 Hz |
ladder resonance | 0.55 |
loudness tap.adsr~ | @attack 8 @decay 300 @sustain -2 @release 350 |
Two moves push it from good to that sound:
- Play the glide. 80 ms of
smoothmeans overlapping note changes swoop; detached ones barely bend. The keyboard articulation is the vibrato. - Lean on the octave voice. Pull voice 3 up to
f(unison) for the hollow reedy register, or leave it atf ÷ 2and drop voice 2's level for the fat fifth-less stack. Theinterp-timed preset morph (store/recall <slot> <ms>) can glide between these voicings mid-phrase — a patch element the hardware never had.
What each ingredient buys
In order — and, per the house rule, cut from the bottom when CPU or taste says so:
- The stack. Three free-running voices at ±cents is most of the sound (the oscillator chapter's argument, with its measurements).
- The ladder. Drive, asymmetry, and the low-
compdroop — the character budget. - The filter contour. The one envelope listeners hear as "the synth's
voice." Its
decayis the most audible 100 ms in the patch. - Glide. Free, iconic, already in the oscillator.
- The loudness contour. Keep it simple; the filter does the talking.
- The analog section.
drift/jitter/imperfectat the chapter's moderate settings — salt, not sauce.
When to leave the recipe
- You want polyphony. This is a monosynth voice;
mc.-wrapping the whole chain gives you many monosynths, and a real polysynth patch wants per-voice envelopes and different discipline. - You want the 303 instead. The couplings that make acid are a
different instrument —
tap.303~refuses to be decoupled, and that refusal is its chapter. - You want clean. Every stage here has an opinion —
tap.svf~andtap.fourpole~are the polite siblings when the patch needs a filter, not a character.
The patches with names on them
The previous recipe built the three-oscillator voice. This one drives it at four records — a blue-eyed-soul hook, the bassline that retired a bass player, a singing art-rock lead, and the one-take modular solo that started it all — and, along the way, answers a fair question: if "Lucky Man" was played on a Moog modular, does the kit need a modular object?
The provenance rule from the part opener applies double here, because gear folklore is a genre of its own. For each patch the chapter says what is documented about the record and what is reconstruction. And the standing disclaimer stands: these settings chase the sound; the hands, the tape, and the mix stay on the record.
Every patch below is a delta against the wiring and tables of Three oscillators into a ladder — build that voice first. Two performance tools recur, so here they are once:
- Vibrato is the oscillator's own now:
@vibrato(depth in cents, so the musical width holds in every register),@vibrato_rate(Hz), and@vibrato_delay(ms) — the onset fades in through that time constant and re-arms on every new note, which is most of what makes a lead "sing." ±10 cents at 5.5 Hz with a few hundred milliseconds of delay is the classic setting. (This chapter's first draft had to print a scaling formula into the Hz-calibrated FM inlet here; that formula became the improvements plan's §2, and §2 became these attributes — the audit worked.) - Sequenced lines:
tap.303.seq~emits pitch as a MIDI-note signal and a gate at 1.0/2.0 —mtof~turns the pitch into Hz for the oscillators' signal inlets, and the gate drivestap.adsr~directly (it opens above 0.5). The Moog voice sequenced this way is the classic synth-line scaffold, slides included.
The Winwood hook — "While You See a Chance" (1980)
What's documented: Winwood played essentially everything on Arc of a Diver himself, synthesizers included; accounts of the rig put Moog monosynths at the center of it. The reconstruction: the opening hook is a brassy, open-filter lead with a fast attack and just enough glide to round the corners — a patch that sits between horn section and organ, which is very much a keyboardist's lead.
Deltas from the lead patch:
| control | setting |
|---|---|
| voices | 1 and 2 only, at f, detune −6 / +6; retire voice 3 |
all smooth | 40 ms |
| ladder | @resonance 0.3 @drive 6 @asym 0.3 |
| filter contour | @attack 5 @decay 500 @sustain -6 @release 400, amount 4500 Hz, base 400 Hz |
| loudness contour | @attack 5 @decay 200 @sustain -2 @release 250 |
The brass illusion is the filter contour's sustain sitting high (−6 dB):
the filter opens and stays open, so the tone holds its brightness through
the note instead of wah-ing. Play the hook in clean detached eighths — the
40 ms of glide only speaks when notes touch.
The bassline that retired a bass player — "Flash Light" (1977)
What's documented, and gloriously so: Bernie Worrell built Parliament's "Flash Light" bassline by stacking Minimoogs — the story is told with the number three attached — playing the line keyboard-style under Bootsy Collins' guitar. This is the patch where the previous chapter's "the stack is most of the sound" rule gets its funk citation.
Deltas from the bass patch:
| control | setting |
|---|---|
| voices | 1 and 2 at f (detune −7 / +7), voice 3 at f ÷ 2, its gain −6 |
all smooth | 35 ms |
| ladder | @resonance 0.6 @drive 12 @asym 0.5 @comp 0.2 |
| filter contour | @attack 1 @decay 150 @sustain -24 @release 150, amount 2200 Hz, base 90 Hz |
| loudness contour | @attack 1 @decay 250 @sustain -6 @release 100 |
The rubber is the filter contour: near-instant attack, short decay, and a
sustain low enough (−24 dB) that every note is a squelch that immediately
ducks. resonance 0.6 puts a vowel on the squelch; drive 12 into the
tanh stages is the fat (the ladder chapter measures 16.5 % THD up there —
that's the point). Play staccato sixteenths with octave pops; let the 35 ms
glide smear only the connected passing notes. If the low end blurs, this is
the one patch where comp earns its raise: 0.2 keeps some droop-era
character while returning enough passband to anchor the root.
The singing lead — "Shine On You Crazy Diamond" (1975)
What's documented: Richard Wright's rig in the Wish You Were Here sessions included a Minimoog, and the singing synth lead lines in "Shine On" are credited to it. The reconstruction: a nearly clean patch — this lead's beauty is restraint, a barely-driven filter, and vibrato that arrives late.
Deltas from the lead patch:
| control | setting |
|---|---|
| voices | 1 and 2 at f, detune −2 / +2 — a shimmer, not a chorus |
all smooth | 15 ms |
| ladder | @resonance 0.15 @drive 3 @asym 0.2 |
| filter contour | @attack 30 @decay 900 @sustain -10 @release 700, amount 3000 Hz, base 250 Hz |
| loudness contour | @attack 8 @decay 300 @sustain -2 @release 500 |
Then spend all your effort on the vibrato: @vibrato 10 @vibrato_rate 5.5 @vibrato_delay 400 — ten cents, arriving late, re-arming on each new
note so held phrase-endings bloom while passing notes stay plain. The patch is deliberately close to the ideal oscillator —
imperfect 0.2, drift at the polite end — because the expressive load is
carried by the hands, and everything the analog section adds here it adds
to sustained exposed notes.
The one-take solo — "Lucky Man" (1970), and the modular question
What's documented: Keith Emerson's solo on "Lucky Man" was played on his Moog modular system and famously kept from an improvised take — one of the first Moog solos on a rock record, and for a generation of listeners the first synthesizer they ever heard. The sound: a huge unison lead whose actual melodic content is mostly portamento — sweeps and dives across octaves, the glide circuit played as the instrument.
Deltas from the lead patch:
| control | setting |
|---|---|
| voices | all three; voice 3 up at f (unison), detune −5 / +4 / +7 |
all smooth | 280 ms |
| ladder | @resonance 0.2 @drive 8 @asym 0.4 |
| filter contour | @attack 10 @decay 800 @sustain -4 @release 600, amount 5000 Hz, base 800 Hz |
| loudness contour | @attack 10 @decay 300 @sustain -1 @release 400 |
At 280 ms of smooth, pitch is a place you travel to: hold a note, strike
one two octaves up, and the voice draws the line between them. That is the
solo. The filter stays essentially open (sustain −4 dB) because the
record's drama is in pitch, not timbre.
So — does the kit need a Moog modular object? No, because you are
holding one. A modular synthesizer is oscillators, filters, envelopes,
and amplifiers with no fixed routing; the panel of patch cords is the
product. In this package the modules are tap.vco~, tap.ladder~,
tap.svf~, tap.adsr~, tap.vca~, tap.noise~, and the sequencer pair —
and Max itself is the patch panel, with the routing freedom no hardwired
monosynth voice (and no single "modular object") could offer. Everything
Emerson's system did on that solo — voices summed to one filter, one
loudness contour, glide on the pitch source — is the previous chapter's
wiring diagram; what the modular added was the freedom to have wired it
otherwise, and that freedom is the patching environment you are already
in. The one genuinely modular idiom worth calling out is the sequenced
line: tap.303.seq~ → mtof~ → the stack, gate → tap.adsr~, is the
Moog-sequencer scaffold of the Berlin school and "I Feel Love"-era disco —
no new object required, slides included.
What separates the four
The instructive part of putting these side by side: the signal chain never changed. What moved:
- The filter contour's
sustain. High and it's brass (Winwood), open and it's drama (Emerson), low and it's rubber (Worrell). One attribute spans the genre map. smooth. 15 ms is articulation, 40 ms is rounding, 280 ms is the melody itself.driveandresonance. The funk patch is the only one leaning hard on both — and it's the one imitating three stacked instruments.- The hands. Delayed vibrato, staccato versus legato, when not to play — the parts of the record the recipe honestly can't ship.
When to leave the recipe
- You want the record's whole arrangement. The hook was never alone: Winwood's is doubled, Worrell's sits under a live band, Wright's floats on tape-delayed guitars. The patch is the voice, not the mix.
- You want polysynth-era sounds. Prophets and Oberheims are a
different architecture — per-voice envelopes on real polyphony — and
imitating them with
mc.stacks of this voice flatters neither. - You want the sequenced-modular genre. Start from the scaffold above, but that recipe deserves its own chapter — it lives in the plan file's backlog with "I Feel Love" written on it.
Move a knob while it loops
The documented origin story of acid house is an instruction manual for this
recipe: in Chicago around 1985–87, Phuture (DJ Pierre, Spanky, Herk) let a
secondhand TB-303 loop a pattern and turned the knobs while it played —
"Acid Tracks" is twelve minutes of that. The lesson generalizes: an acid
line is not a melody with a sound; it is a loop plus a hand. The
pattern's job is to give the couplings something to chew on — accents for
the bloom, slides for the vowels — and the performance is cutoff,
resonance, and envmod moving in real time.
Everything measured here is borrowed from the acid machine
chapter and its notebooks
(tb303.ipynb,
step_seq.ipynb).
The scaffold
phasor~ (BPM/240) ──▶ tap.303.seq~ ──pitch──▶ tap.303~ ──▶ out
└──gate───▶ (right inlet)
One bar of 16 steps per phasor cycle (BPM ÷ 240, the drum scaffold's math);
~125 BPM → phasor~ 0.5208. The sequencer's pitch and gate outlets are the
voice's own contract — accents ride the gate at 2.0, slides are pitch
changes under a held gate, so the voice's ~60 ms RC does the glide.
A line to start from
Program per step (step <n> <pitch> [accent] [slide], rest <n>) or per
lane. A serviceable opener in A — and, as everywhere in this part, a
starting point, not a transcription:
step: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
pitch: 33 33 45 33 33 36 33 31 33 33 45 47 33 33 31 36
gate: x x x x . x x x x x x x . x x x
accent: A . . A . . A . . A . . . . A .
slide: . . . . . . . S . . . S . . . .
pitches 33 33 45 33 33 36 33 31 33 33 45 47 33 33 31 36
gates 1 1 1 1 0 1 1 1 1 1 1 1 0 1 1 1
accents 1 0 0 1 0 0 1 0 0 1 0 0 0 0 1 0
slides 0 0 0 0 0 0 0 1 0 0 0 1 0 0 0 0
The ingredients that make it acid rather than bass: the octave jumps (33 → 45), at least one slide into a note (the flag sits on the target step), rests that let the filter close, and accents placed where the groove leans — not where the melody peaks.
The voice
recall 1 is the factory "squelch" and a fine start. Explicitly:
@waveform saw @cutoff 500 @resonance 0.9 @envmod 0.7 @decay 300 @accent 0.8
Then the moves, in the order a set builds:
- Ride
cutoff. 300 → 2000 Hz over eight bars and back. This is the genre. Remember the modeledenvmodlaw: 2/3 of the envelope's sweep sits above the knob, 1/3 below, and the resting point shifts as you turn it — the knobs interact like the hardware because the interaction is modeled. - Stack the accents. Runs of accented notes at high
resonanceare the wow: the C13 capacitor doesn't fully discharge between close accents, and the measured cutoff-peak bloom across a run is ×1.94. Put three accents in a row somewhere and listen to the third one open. - Raise
resonanceinto the break. Past 1.0 is the documented bend territory — a stock 303 never self-oscillates, and neither does this filter until you push it there deliberately. waveform squarefor the hollow verse, saw for the drop.vca warmthickens exactly where the hardware does — measured 5.4 % difference signal on quiet notes, 11.5 % on hot accents.- Transpose, don't re-program:
transpose -12on the sequencer for the sub-drop,+5/+7for the question-answer sections. It shifts live, like the hardware's transpose mode without the mode.
After the voice
Acid techno's other instrument is the distortion pedal: tap.overdrive~
after the voice, driven hard, is the documented lineage (a 303 into a
screaming feedback overdrive is half the harder end of the genre). Keep
mute in reach on the sequencer for breakdowns — it drops the gate but the
clock keeps running, so the line re-enters exactly in place.
When to leave the recipe
- You're programming melodies. If the line only sounds right without
slides or accents, it isn't an acid line yet — or it wants the generic
bass rig (
tap.vco~+tap.svf~+tap.adsr~) instead of this voice's refusals. - You want the filter alone —
tap.diode~gives you the 303's ladder on any source, squelch and all, without the biography. - You want hands-free evolution. The 303 rewards a hand on a knob; if
the patch must run itself, store extremes in the voice's preset slots and
ride timed
recallmorphs instead — a different instrument, honestly.
The ostinato machine
Two documented lineages share one patch. The Berlin school — Tangerine Dream's Phaedra (1974) above all — put a Moog modular's step sequencer on stage and let a filtered ostinato run for twenty minutes while hands moved the cutoff. Three years later Giorgio Moroder and Donna Summer's "I Feel Love" built an entire hit from a sequenced Moog modular bassline. The Recipes part's Moog chapter argued you already own the modular — Max is the patch panel; this recipe is that argument cashed in: the sequencer pair driving the three-oscillator voice.
The scaffold
phasor~ (BPM/240) ──▶ tap.303.seq~ ──pitch──▶ mtof~ ──▶ slide~ ─┬─▶ tap.vco~ ──┐
│ └─▶ tap.vco~ ──┼─▶ *~ ─▶ tap.ladder~ ─▶ tap.vca~
└──gate──┬─▶ tap.adsr~ (filter) ─▶ *~ amount ─▶ +~ base ──▶ ▲ (cutoff)
└─▶ tap.adsr~ (loudness) ────────────────────────────▶ ▲ (gain)
tap.303.seq~'s pitch outlet is a MIDI-note signal;mtof~turns it into Hz for the oscillators' signal inlets.- Its gate outlet (1.0 plain, 2.0 accented) drives both
tap.adsr~contours directly — the envelope opens above 0.5. - One honest wrinkle:
tap.vco~'s frequency signal inlet bypasses thesmoothramp by design ("you are the smoothing") — so sequenced pitch steps land as hard steps, and slide flags in the pattern won't glide on their own. Put a one-pole slew (Max'sslide~, orrampsmooth~) betweenmtof~and the oscillators; the 303's ~60 ms RC is the reference feel. Short slew = articulation, long slew = the Berlin swoop. - The oscillator stack, ladder voicing, and envelope tables come from Three oscillators into a ladder — the bass patch is the right starting point. One voice instead of three is period-correct for the sequenced genre and cheaper; add the stack when the line is the whole arrangement.
The line
The genre's cell is small and the sequencer's phase math does the rest
(one bar per phasor cycle; a length 8 row divides it into eighths —
polymeter as arithmetic, per the sequencer chapter).
The octave bounce, "I Feel Love"-school — length 8, every step gated:
step: 1 2 3 4 5 6 7 8
pitch: 33 45 33 45 33 45 33 45
pitches 33 45 33 45 33 45 33 45
gates 1 1 1 1 1 1 1 1
The Berlin cell — length 16, a contour that repeats but doesn't resolve:
pitches 33 33 40 36 33 43 36 40 33 33 40 36 31 43 36 38
gates 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
Then the two moves that carry twenty minutes:
- Transpose, sparsely.
transpose 0 → 3 → 5 → 0at phrase boundaries is the harmonic language — one message, and the armed pattern semantics keep everything on the grid. - Ride the filter, slowly. The loudness contour stays short and
percussive; your hand (or a very slow LFO into the
+~ base) opens the ladder over minutes, not bars. The ostinato doesn't change; the light on it does.
Settings that read as the genre
| control | setting |
|---|---|
loudness tap.adsr~ | @attack 2 @decay 180 @sustain -12 @release 120 |
filter tap.adsr~ | @attack 2 @decay 160 @sustain -20 @release 160, amount 1800 Hz, base 150 Hz |
| ladder | @mode lp24 @resonance 0.3 @drive 6 |
slew (slide~) | short; raise it only for deliberate swoops |
| seq | @swing 0 — the genre is a grid, and the delay does the humanizing |
Two period tricks worth their lines: pan alternate notes (a length 8 row
of accents driving tap.pan~ recreates the famous ping-pong doubling), and
put an eighth-note tap.delay~ after the voice (@feedback 40 @mix 30) —
the echo, not the sequencer, is where these records' motion lives. Since
its rebuild the delay interpolates (Hermite) and regenerates through a
DC-blocked loop; @interp 0 remains the bit-faithful legacy mode.
Glue: an 808 closed-hat row in 16ths from the drum scaffold, mixed low.
Accents land in this scaffold too: turn up tap.adsr~'s velocity
sensitivity and the sequencer's 2.0-amplitude accented gates hit the
envelopes harder — the loudness contour for punch, the filter contour for
the quack, or both.
When to leave the recipe
- You want the 303's couplings — slides that bloom, accents that squelch. That's the acid recipe; this scaffold trades the couplings for a filter and envelopes you choose.
- You want generative movement. This sequencer is deliberately
deterministic; probability and ratchets are future emitters, and
randomness belongs to objects that own a
seed. - The line wants to be a song. Sixty-four steps is the ceiling; past
that you're composing, and a piano roll is kinder than sixty-four
stepmessages.
The robot on the radio
The vocoder chapter closes on "the casting is everything," and this recipe is the casting call. The sound has a documented pedigree — Bell Labs speech-compression research became, in musicians' hands, Kraftwerk's robot choirs and ELO's talking skies, and the machine has never left the radio since — but the records disagree on gear and agree on craft: a bright, busy carrier, an articulate modulator, and somebody enunciating like they mean it. All three are patching decisions.
Everything structural below is pinned by the kernel's tests and explained in the vocoder chapter: 24 bands, 50 Hz–12 kHz, everything you hear is carrier.
The carrier, built properly
The eternal failure is a dull carrier — high bands with nothing in them, consonants gone. Build it in three layers:
tap.vco~ (saw, f) ──┐
tap.vco~ (saw, f, +7 c) ──┼─▶ +~ ──▶ tap.vocoder~ right inlet
tap.noise~ (white) ─ *~ 0.1┘
- Two saws, a few cents apart (
@shape 2,detune±4–7): harmonics to the top of the range, and the beating keeps long vowels alive. Use the Moog recipe's stack values; skip the octave-down voice — vocoded speech reads clearest with the energy above the fundamental. - The s and t budget.
@sibilance 0.3is the built-in version — a seeded noise source in the top bands' carrier, gated by the modulator's high-band envelopes, arriving exactly when consonants do. The manual alternative (a tenth oftap.noise~summed into the carrier) remains the craftier option when you want to choose the noise color yourself. - Pitch is the performance. The vocoder never changes the carrier's
pitch, so the carrier's notes are the melody. Held chords (an
mc.stack of carriers) make the robot a choir; a single line makes it a lead vocalist.
The modulator, cast against type
Articulation beats fidelity — the chapter's measured point is that band envelopes carry everything, so contrast between bands is what you feed it. A cheap dynamic mic is fine; compression helps; and over-enunciating helps more than any knob. Keep the modulator out of the mix — the machine uses it, nobody should hear it.
The three settings
| patch | q | response_interval | the craft |
|---|---|---|---|
| the talking synth | 20 (default) | 30 | speak in rhythm; consonants land like drum hits |
| the choir | 10–15 | 250 | sing sustained vowels; the carrier chord is the harmony |
| rhythm transfer | 25–40 | 10–20 | drum loop as modulator; any sustained pad as carrier |
q trades crispness against smoothness (narrow separates consonants,
wide blends vowels); response_interval is the mouth's speed — attack and
release in one knob. gain is linear makeup, and you will need some: a
band-multiplied signal lands quieter than either input.
Two wiring facts that account for most dead patches: the modulator is the left inlet (a synth weakly filtered by your voice means the cables are backwards), and a silent carrier is silence no matter how loudly you speak — pinned by test, and the fastest debugging question in vocoding.
The songbook
The famous "vocoder songs" are the best syllabus for the craft — partly because several of them aren't vocoders, and knowing which is which teaches more than any preset. Provenance below follows the part's rule: documented where it's documented, labeled reconstruction where it isn't.
"In the Air Tonight" (1981) — the ghost choir
What's documented: Phil Collins ran the verse vocal through a Roland
VP-330 — a soft vocoder, voiced like a string machine, mixed under a
nearly whispered dry vocal. The reconstruction: this is the anti-robot
patch. Carrier: two saws, detune ±4, imperfect 0.3, no noise
layer — sibilance is what you don't want here — through tap.svf~
(@type lowpass @frequency 4000) to take the glass off. Vocoder:
q 8–12, response_interval 120 — wide bands and a slow mouth blur the
consonants into breath. Mix the vocoded return under the dry voice, a
shadow rather than a double. The dry whisper carries the words; the
vocoder carries the dread.
"Mr. Blue Sky" / "Mr. Roboto" / "Intergalactic" — the front-and-center robot
ELO (EMS vocoders, documented), Styx, and the Beastie Boys are the
talking-synth patch played as a lead: bright carrier, crisp bands
(q 20–30), fast mouth (response_interval 15–30), and the melody in
the carrier's held notes while the words ride the modulator. Kraftwerk —
the genre's founders, on custom and commercial hardware across the
years — sit here too, usually with a single unison line rather than
chords: the robot speaks in monophony. Enunciate. Then enunciate more.
One label to keep straight: Zapp, Roger Troutman, and the P-Funk talkbox records are not vocoders — a talkbox pipes the carrier into the performer's actual mouth and the room mic hears real articulation. Chasing that sound with this object gets you a cousin, not the thing.
"Hide and Seek" (2005) — the one that isn't a vocoder
What's documented: Imogen Heap sang into a harmonizer (the DigiTech Vocalist lineage), a keyboard choosing the chord — so every sound on the record is her actual voice, pitch-shifted into harmony, breath and formants intact. That's why it doesn't sound like a robot; there is no carrier. Three routes, honestly ranked:
- The right tool —
tap.harmony~. This record's mechanism is exactly what the object does: formant-preserving voices holding a chord over the aligned dry voice. Its recipe — with the Bon Iver patches that extend the lineage — is A choir of one. (This object exists because this section's first draft had to work around its absence; the audit worked.) - The manual fallback — a shifter stack. Voice into parallel
tap.shift~objects at chord intervals (tap.semitone2ratiofeeds their ratio inlets). Keep the voicing within ±7 st — plain granular shifting moves formants with the pitch, and wide intervals go chipmunk where the formant-corrected routes don't. - The vibe — the choir patch above. Speak-sing into the choir row's
settings with an
mc.carrier holding the chords. It will sound like a vocoder doing Imogen Heap, which is its own valid sound — just don't mistake it for the record's mechanism.
The plugin-era default — an Orange-school carrier
The late-90s software vocoders (the Orange Vocoder the most loved of them) changed the default sound of the effect: where hardware vocoders leaned on whatever synth was nearby, the plugins shipped with a built-in, very bright virtual-analog carrier — so "the plugin sound" is really a carrier voicing: wide, glossy, present. One honest line first: that plugin is still a shipping commercial product, and the house provenance rule applies — nothing here reverse-engineers it. What follows is our bright VA carrier in that school, built from this package's own oscillator:
tap.vco~ (saw, f, detune -6, seed 11) ──┐
tap.vco~ (saw, f, detune +6, seed 22) ──┼─▶ +~ ─▶ tap.svf~ (highshelf) ─▶ carrier
tap.vco~ (saw, f+12, gain -6, seed 33) ──┤
tap.noise~ (white) ─ *~ 0.08 ─────────────┘
All three oscillators @shape 2 @imperfect 0.2 @jitter 2; the octave-up
voice adds the gloss the era is remembered for; tap.svf~ @type highshelf @frequency 6000 @gain 4 is the sheen. Vocoder settings:
q 25, response_interval 20. Play the carrier in fifths and octaves
rather than full triads — the brightness supplies the width, and triads
in a bright carrier smear the consonant bands.
When to leave the recipe
- You want tuned speech, not a played carrier — the corrector
(
tap.tune~) moves the voice itself; the vocoder wears the voice over something else. Different identity theft. - You want formant-shifted or gender-shifted voice — that's spectral
surgery, not band gating; the corridor starts at
tap.spectra~. - You want intelligibility above all. Twenty-four analog-style bands are a voice, not a spectrograph; if every syllable must survive, dry speech mixed under the vocoded double is the radio trick that always works.
A choir of one
This chapter exists because this book's own audit demanded it. The vocoder
songbook had to label "Hide and Seek" honestly — a harmonizer, not a
vocoder: pitch-shifted copies of the actual voice, formants intact, no
carrier anywhere — and the package had no object for that mechanism. Now it
does. tap.harmony~ holds up to four formant-preserving voices at
intervals you set in semitones, over a dry path the kernel delays into
alignment so chords land as chords. This recipe is how to sing through it,
and its worked examples are the modern masters of the effect: Bon Iver.
The claims behind the object are measured in the executed
verification notebook
and pinned in the kernel's test battery (tests/harmonizer_test.cpp):
across two octaves of voicings every interval lands within 0.04 cents
of its equal-tempered target under the DspTap yin oracle; the dry path is
sample-aligned with the voices to a 3.7×10⁻⁸ residual — why chords land
as chords, not flams; a synthetic formant bump stays near home only when
formant is on (band centroid 1058 → 1154 Hz on a +7 shift, versus 1439
riding the full ratio with it off); and interval glides walk the pitch
through the middle instead of jumping. The engine per voice is the DspTap
phase vocoder — the same peak-locked shifting and LPC formant machinery
the pitch machine chapter derives.
The instrument
voice ──▶ tap.tune~ (@speed 0, key of the song) ──▶ tap.harmony~ ──▶ out
The corrector upstream is optional but it is the modern sound: hard-snap the lead first and every harmonizer voice inherits the quantized pitch, so the stack locks like a keyboard instrument instead of drifting like a choir. Skip it and the stack breathes with your intonation — older, warmer, more Crosby-Stills than Vernon.
Three controls do the character work:
formanton is the entire point: an octave-down voice stays you, bigger. Off is the chipmunk-chorus bend — useful, but it stops being a choir.chordis the performance surface:chord -12 3 7sets three intervals and silences the fourth voice in one message. Wire a Max chord-to-intervals mapping (played notes minus the sung root) and the keyboard chooses the harmony live — the rig the credits of the records below describe.glideat the 10 ms default snaps chord changes; at 300–500 ms the stack slides between chords, which no group of human singers can do and is worth featuring, not hiding.
One honest number: latency is one FFT frame (fftsize, default 1024
samples ≈ 21 ms at 48 kHz), dry path included. For live monitoring that
is audible as a slight remove — performers adapt in minutes, but mix the
monitor wet so they hear the instrument, not the delay ghost.
"Woods" (2009) — the stacked chapel
What's documented: Justin Vernon built "Woods" from many overdubbed a cappella takes, each hard-tuned — a chapel of his own voice, later the foundation of Kanye West's "Lost in the World." The record's mechanism is overdubs, and the recipe respects that:
- The live approximation:
tap.tune~ @speed 0→tap.harmony~withchord 3 7 12,@dry 1, all through a long dark reverb (the wash settings work). One pass, four-voice chapel. - The faithful version: record takes — sing each chord tone through the
corrector alone, layer them, and use
tap.harmony~per take only to widen (chord 12at@level1 0.4). Stacked takes decorrelate the way overdubs do; one harmonizer pass, however good, is one performance. The difference is the difference between a choir and a string patch.
"715 - CRΞΞKS" (2016) — the Messina
What's documented: the 22, A Million credits name "the Messina," the rig Chris Messina and Vernon built to pitch-stack his live voice into chords (the Prismizer-school effect associated with Francis and the Lights, and heard on Chance the Rapper's gospel records). "715 - CRΞΞKS" is that instrument a cappella: every sound is the processed voice.
voice ─▶ tap.tune~ (@speed 0) ─▶ tap.harmony~ @dry 1 @formant 1 @glide 10
chord -12 3 7 (verse color)
chord -12 4 7 (the lift)
chord -5 3 10 (the ache)
@dry 1— the lead lives inside the stack, equal citizen, exactly what makes the sound read as one multiplied person rather than lead-plus-backers.- Chords change per phrase, not per note: bind each
chordlist to a key or a pedal and play the harmony like slow organ stops. - The low voice carries the weight:
-12under a falsetto lead is the record's signature register trick. Keep it at full level; thin the upper voices (@level3 0.7) when the stack gets glassy. - No reverb, or almost none — the record's intimacy is the dry stack right against the microphone. Resist the wash this once.
The craft notes
- Feed it one voice. The formant model and the intervals both assume monophonic input — the kernel's header says so, and a strummed guitar through a "choir" proves it right within two bars.
- Mind the sum. Dry plus four unity voices is five voices;
tap.limi~or a*~ 0.5after the object is the standing advice. - Close voicings beat wide ones. ±12 is the working span; the engine's contract runs to ±24 and the top octave of that range is a sound effect, not a singer.
- The corrector's
speedis the era dial. 0 ms is 2016; 40 ms is 1970s session stack; bypassed is a folk group.
When to leave the recipe
- You want the robot. No carrier here, no bands — that's the vocoder, and the two chapters together are the voice-processing fork in the road: wear the voice over a synth, or multiply the voice itself.
- You want real ensemble. Overdub real takes; the "Woods" section's faithful version is the honest ceiling of one person's choir.
- You want harmony that follows chords you sing. The object holds intervals; it doesn't do music theory. The keyboard (or your patch's chord logic) is the brain — which is exactly how the famous rigs work.
The staircase and the wash
Shimmer has a documented birthplace: Brian Eno and Daniel Lanois in the early eighties, feeding a pitch shifter and a reverb into each other until a guitar came out sounding like weather. The pitchaccum chapter tells the half of the story that lives inside one object — the transposer-delay loop where every pass climbs again — and its recipes sketch the pairing. This recipe is the whole patch: the spiral, the wash, and the mix decisions that keep ten seconds of accumulated fifths from eating a track.
Measured claims are borrowed from the spiral staircase (the +7-becomes-+14 accumulation, the constant-sum grain envelopes, the 0.99 feedback cap) and borrowed rooms.
The chain
source ──▶ tap.pitchaccum~ ──▶ reverb (tap.verb~ or tap.convolve~) ──▶ return
└────────────────── dry path ───────────────────────────────────────▶ out
Run it as a send: the source stays dry and full-size in the mix, and the
shimmer return comes up underneath it like backlight. On the send,
tap.pitchaccum~ at mix 100 (its own dry path stays home) and the
reverb wet-only.
The spiral
| control | setting |
|---|---|
trans1 / delay1 / fb1 / gain1 | +12 st / 400 ms / 75 / 50 |
trans2 / delay2 / fb2 / gain2 | +7 st / 650 ms / 60 / 50 |
xfade | 60 — smooth flanks, soft attacks |
modfreq / moddepth / modphase | 0.3 Hz / 0.1 st / 90° |
follow | off for chords and pads; on for monophonic lines |
The two shadows are doing different jobs: the octave climbs politely
(+12, +24, +36 — always consonant), while the fifth rotates the harmony
(+7, +14, +21 — a fifth, then a ninth, then a #11) and is where the
Eno-school mystery comes from. Pull fb2 down toward 40 when the source
is already harmonically rich; push fb1 toward 90 for the endless
version — the loop is capped and DC-blocked, so "too long" is an
aesthetic problem, not a stability one. The touch of modulation
(moddepth 0.1, with modphase 90 breathing the shadows against each
other) keeps a long spiral from sounding cloned — depth stays subtle or
the climb turns seasick.
The wash
Either reverb works; they fail differently:
tap.verb~(the designed tail):@mix 100 @decay 8 @damping 4000 @lowpass 8000 @delay 60 @modfreq 0.2 @moddepth 0.3. The damping matters more than the length — shimmer's accumulated highs need somewhere soft to land, and 4 kHz of loop damping is the difference between glow and glass dust.tap.convolve~(the borrowed room): a long church at@mix 100 @predelay 20, and pick the IR by its top end — audition the tail alone and reject anything that rings metallic up high, because the spiral will find it. (The field guide to rooms has the audition drill.)
Order matters and is worth an experiment: spiral → reverb (above) washes the staircase — the classic. Reverb → spiral transposes the wash itself and is wilder and less controllable; the historical chains did both, depending on the record.
Variants
- The descent:
trans1 -5,trans2 -12, long delays, feedback ~50 — the staircase into the basement. Darker damping (2–3 kHz); the low accumulation muddies fast, so shorter reverb. - The micro-halo:
trans1 +0.15,trans2 -0.15, delays 60/90 ms, feedback ~50,xfadewide, modest reverb — no spiral at all, an expensive-sounding widener that flatters pads. - The gesture: store the halo in slot 1 and the full +12/+7 spiral in
slot 2, then
recall 2 8000as the chorus lands — the morph engine glides every parameter, and the bloom is the production moment.
When to leave the recipe
- The mix is dense. Shimmer is backlight; on a busy arrangement it reads as mud. It earns its keep on sparse sources — one guitar, one voice, one held pad.
- You want rhythmic echoes climbing in pitch. The delays here serve
the loop, not the grid; that patch is a tempo-synced delay into
tap.shift~, built by hand. - You want the pitch to stay put. Then it's just reverb — go straight to the field guide.
Sixteenths into a listening filter
The envelope filter earned its place in funk on documented records — the Mu-Tron-era clavinets and basses of the seventies, Stevie Wonder's "Higher Ground" chief among them — and the autowah chapter is honest about what this object is instead: a model of the Snow White AutoWah, a different, throatier circuit. You are not summoning a Mu-Tron; you are plugging into a very good pedal that listens the same way. The funk is in what it listens to — which makes this the one recipe where the settings table is half the story and your right hand is the other half.
Measured behavior cited below (the sweep law exact to the design, the RC release, the 250 Hz → ~2.5 kHz hardware span) lives in the pedal that listens and its validation notebook.
How to think about the knobs
Two of them are calibration, one is the personality:
sensitivitymatches the pedal to your source's level and your touch. Tune it so your normal hits open the filter halfway and your hard hits open it fully — the tanh knee compresses beyond that instead of slamming. Too high and everything pins; too low and the filter ignores you.biasandrangeset where the sweep lives: resting frequency and octaves above it. The defaults (250 Hz, 3.3 octaves) are the hardware.decayis the personality: how fast the filter falls back. Tens of ms is a wah articulation on every note — the funk setting. Hundreds is a swell that rides phrasing.
The patches
| patch | settings |
|---|---|
| the clav chop | @sensitivity 3 @attack 2 @decay 80 @bias 250 @range 3.3 @resonance 0.7 @mix 100 |
| the bass quack | recall 2, then @decay 150 @resonance 0.6 |
| the slow swell | recall 3, or @decay 900 @range 2.5 on pads |
| the cocked wah | recall 4 — sensitivity at −60 is the envelope off; park bias at 800–1200 Hz |
- The clav chop wants sixteenth-note playing with deliberate dynamic
contrast — the filter turns your accents into vowels.
mode 1(bandpass, the circuit's other tap) is quackier and noticeably quieter; make it up withgain. - The bass quack starts from the factory bass voicing (slot 2 — lower bias, tighter range, the GB pedal's instrument switch as a preset). Fingers, not pick, and let notes ring — the release is a real RC discharge (measured: a pure exponential, σ = 0.004) and it sounds like circuitry when you leave it room.
- The cocked wah is the secret mode: a fixed resonant filter with
biasas a manual sweep — the parked-pedal midrange honk, and slot 4 ships it. direction 1sweeps down from bias — the extension the pedal never had; reverse-envelope funk on a clean chop is startling.
The two patch points nobody uses enough
- The sidechain (right inlet): one sound wahs another. Kick →
sidechain, pad → filter is the classic; a
tap.808.seq~row (through@pulsewidened impulses) makes the filter sequenced while the pad sustains — an envelope filter with a drummer's timing. - The envelope outlet (right outlet, 0..1 signal): the detector as a
free modulation source. Scale it into
tap.vco~'s FM inlet, atap.vca~gain, or a second filter — one performance, many destinations. (Inbypassthe outlet goes to zero, so tap it from a live instance.)
When to leave the recipe
- Your source has no dynamics. A static pad through an auto-wah is a
static filter — feed the sidechain something rhythmic, or use
tap.svf~with an LFO and own the motion yourself. - You want the filter on a knob. That's the cocked wah until you want
morphing responses — then
tap.svf~'smorphis the tool. - You want the exact Mu-Tron quack. Raise
resonance, trymode 1, and know the chapter's warning stands: you're modding a Snow White. The hardware A/B pass — the notebook cell waiting for the real pedal — will tell us precisely how far the model is from its own hardware, not from someone else's.
A field guide to rooms
The convolution chapter makes one promise that changes how you shop:
tap.convolve~ is exact — measured to 10⁻¹² against direct convolution —
so the engine contributes no character at all. Everything the effect sounds
like is the impulse response you load. That turns "how do I get a good
reverb?" into "how do I find, judge, and place a good room?" — a curation
problem, and this recipe is the field guide.
Companion material: the convolution chapter and its verification notebook; every measured number cited below lives in one of them.
The shopping list
An IR is a recording of a space answering a click, and the internet holds decades of them — university acoustics archives, church-recording projects, hardware-unit captures released by their communities. What to bring home, by job:
- A church or concert hall (2–5 s). The default "make it beautiful" space. Long tails flatter sustained, sparse material and drown busy mixes — the classic trade.
- A plate. Not a room at all — a steel sheet's dense, fast-building wash. The vocal reverb of half the records you know; sits in a mix better than any hall because it has no early-reflection "walls" to argue with the stereo image.
- A spring. The lo-fi twang of amp reverb; gloriously wrong on drums.
- A small real room (0.3–0.8 s). The most useful and least glamorous purchase: drums and guitars recorded dry come alive with a believable space that reads as "a room," not "an effect."
- Not a room. The chapter's point stands in practice: any filter you can record is loadable. A guitar body IR makes a piezo pickup sound like wood; a vowel is a formant filter; a single click is a delay.
Prefer 4-channel captures when offered: the engine runs the full true-stereo matrix (LL/LR/RL/RR), and the cross-feed paths are where "being in the room" lives — measured in the notebook at exact path gains with zero leakage. A 2-channel IR runs as dual mono (no cross-feed); a mono IR is the same room on both sides.
Judging a room in sixty seconds
Load it into the buffer~, then:
- Send a click through and listen to the tail alone (
mixfully wet). You are auditioning the IR itself — the engine adds nothing. A good tail decays smoothly darker; a flutter or a metallic ring here will be on everything you send. - Check the onset. Silence before the direct sound is pre-delay baked
into the capture — trim it in an editor or accept it, but know it's
there, because it stacks with the
predelayyou set and the engine's ownblocksizesamples of latency. - A/B at matched loudness.
normalize 1is on by default and is energy-based, so a quiet cathedral capture and a hot plate land at comparable levels — judge the room, not the gain staging.
Placing the room in a patch
- Send, don't insert. One
tap.convolve~fed by a send bus serves the whole patch, glues sources into one space, and keeps the option of riding the send. Keepmix 100(wet-only) on a send; usemixas an insert dry/wet only on a single source. predelaybefore you EQ. 10–30 ms separates the dry attack from the wash and buys clarity for free — the chapter's advice, and the first knob to reach for when a reverb "swallows" a vocal.blocksizeby role. Live input through the reverb: 64–128 (1.3–2.7 ms at 48 kHz — the measured cost is exactlyblocksizesamples). Mix-bus send: 512–2048, the CPU-cheap end, where the latency reads as a little extra pre-delay you set once and forget.- Swap rooms as a performance move. IR swaps are atomic and click-free
(measured RMS across the swap: 21.9 → 22.1) — load verse-room and
chorus-hall into two
buffer~objects and rebind withset <buffer-name>on the downbeat. (The buffer is the only way in: the object binds thebuffer~named by its first argument, and re-loading a file into that buffer re-transforms the IR automatically.)
When to leave the recipe
- You want to design the tail — decay and damping knobs, modulation,
gated endings. A static IR can't;
tap.verb~is the algorithmic sibling built for exactly that. - You want shimmer. The wash is only half of it — the spiral half is
tap.pitchaccum~, and that pairing has its own recipe in this part. - You want zero latency.
blocksizesamples, full stop; at 64 that's small, not zero.
Chords with no keyboard
The comb-bank chapter ends on "strings, chords, drones, and gestures; no
guitar required" — this recipe supplies the chords. tap.5comb~'s five
voices tune in Hz (freq1..5), which means voicings are numbers you can
keep, trade, and morph between; below is a small book of them, plus the
excitation and morph craft that turns a filter bank into an instrument.
The mechanics cited here — ring time on a log map (20 ms–100 s), Hermite
tuning, warp's stretched partials, phase's midpoint pluck — are
measured and explained in five strings, no guitar and
its machine chapter.
Voicings to keep
Tunings in Hz; MIDI equivalents in parentheses for orientation. The
notes message tunes a voicing in one gesture — up to five MIDI note
numbers, fractional allowed, so just-intonation intervals land exactly
(notes 45 52 57 60.86 64 is the major glow with its true 5/4 third at
275 Hz) — and the Hz attributes remain for exact ratios like the bell
plate.
| voicing | freq1..5 | character |
|---|---|---|
| the factory chord | 80 / 120 / 160 / 200 / 102 | the legacy preset: a root-fifth-octave stack with a rub (102 against 80) |
| the open fifth | 55 / 110 / 165 / 220 / 330 (A1, A2, E3, A3, E4) | power-chord drone; nothing to clash with any source |
| the major glow | 110 / 165 / 220 / 275 / 330 (A2, E3, A3, ~C#4, E4) | just-intonation major: 275 is a pure 5/4 third, warmer than 12-TET's 277.2 |
| the dark cluster | 65.4 / 77.8 / 98 / 130.8 / 196 (C2, D#2, G2, C3, G3) | minor with a low rub; film-cue territory |
| the bell plate | 210 / 297 / 420 / 594 / 841 | non-octave (√2 ratios): inharmonic, gong-ish before warp even arrives |
Masters make voicings performable: freq (0..2) transposes the whole
bank proportionally — chords stay chords under the glide — and res/lp
scale ring and brightness bank-wide.
Ringing them
- Drone:
res1..5 85,lptoward 5000. Feed it anything quiet and sustained — pink noise at low level, a field recording, your room tone. Atres≈ 100 the bank sustains essentially forever; the input stops being audio and becomes bowing pressure. - Pluck:
res1..5around 60–70 and excite with clicks or a sparsetap.808.rim~(@model claves) pattern — every tick strums the chord. Drums work; speech works eerily well (the chapter's "resonator chord"). - Strings, stiffened:
warp 40stretches the upper partials sharp — piano-ish, then bell-ish — while the compensated main tap keeps the pitch put. Pair withlpnear 3000 for the felt-hammer version. - The midpoint pluck:
phase 100cancels the even harmonics — the hollow, clarinet-adjacent voicing of a string plucked exactly at its middle. On the bell plate tuning it turns purely ceremonial.
Watch the sum: five ringing combs stack like five strings. Ride gain
down as res goes up, and tap.limi~ on the output is cheap insurance
for the res 100 lifestyle.
The gesture
The bank's real instrument is the morph engine. Store the major glow in
slot 1 and the dark cluster in slot 2 — then recall 2 8000 and every
frequency, ring time, and damping glides for eight seconds through
tunings you never chose, Hermite interpolation keeping the sweep
continuous instead of zippered. The chapter's advice stands: automate
nothing else. One long morph over a static source is a complete piece of
sound design; grabbing a single fader mid-morph overrides just that
parameter, which is the escape hatch when the in-between territory finds
something worth keeping.
When to leave the recipe
- One resonance, surgically placed:
tap.comb~is the single unit, ortap.svf~ @type bellwhen you want EQ, not a string. - You want echoes. Combs long enough to hear as repeats are delays
wearing a costume —
tap.delay~/tap.multitap~are the honest tools. - You want more than five strings. The count is fixed;
mc.wrapping the whole bank gives you choirs of banks, at which point you are building a sympathetic-string instrument and should budget CPU like it.
The part that comes apart, on tape
Three objects and one posture: hands on the controls while it runs. The stutter takes a phrase apart, the tape echo smears the pieces, and the fuzz decides how hard the whole thing is being pushed. None of the three has a "right" setting, which is the point of the part of the book they live in.
This recipe is a rig, not a record. What is documented about the Kid A-era working method is that a laptop running Max sat in the signal path and got played — the objects here are informed by what that rig was for, not by anyone's patch. Nothing below is claimed to be a reconstruction of a specific track.
Measured claims are borrowed from four heads and a motor, the part that comes apart and the dirt with two stages.
The chain
source ──▶ tap.fuzz~ ──▶ tap.stammer~ ──▶ tap.tapecho~ ──▶ out
Order matters, and this order is the useful one:
- Fuzz first, because a stutter of a distorted signal is a stutter; a distortion of a stuttered signal turns every slice edge into a transient the clipper amplifies.
- Echo last, because the echo is the only object here that is supposed to blur. Put it before the stutter and the stutter slices the blur, which sounds like a mistake rather than a decision.
The dirt
Keep it low. Two saturators in series get muddy fast, and the echo has its
own drive.
| control | setting |
|---|---|
gain | 0.35 |
edge | 0.3 |
asymmetry | 0.2 |
bass / treble / contrast | 0. / 0.1 / 0.3 |
oversample | 4 |
level | to taste, usually negative |
contrast is the scoop, and a scooped source stutters better than a
mid-heavy one — the slices stop fighting the vocal or the guitar they came
from. Leave oversample at 4 and resist the urge to save the cycles at 2:
a hard edge on bright material is exactly the case where one doubling
stops being enough, and 2 measures badly above about 6 kHz.
The stutter
| control | setting | what it does |
|---|---|---|
step | 60. ms | the grid |
divisions | 1 | |
density | 0.3 | how often it grabs |
repeats | 4 | how long it holds |
reverse | 0.2 | chance a repeat plays backwards |
jump | 250. ms | how far back a slice may reach |
fade | 2. ms | the flanks |
seed | any integer | |
mix | 100 |
The one thing to internalize: repeats is the hold, density is the
grab. Occupancy — how much of the timeline has a slice in flight —
measures 41 % at density 0.3 / repeats 1 and 90 % at density 0.9 /
repeats 1, but density 0.3 with repeats 6 already sits at 76 %. If the part
feels too busy, pull repeats before you touch density; you will keep
the sparseness of the entrances while shortening what each one does.
seed is a contract, not a flavour: the same seed and the same moves give
a bit-identical render, and two instances on two tracks decorrelate by seed
alone. Different seeds change 89 % of samples, so it is a real dice roll,
not a phase tweak.
At density 0. the object is a bitwise bypass — worth knowing, because it
means you can automate density to zero and get the dry signal back exactly,
with no crossfade artifact to work around.
The material contract. This object flatters a played phrase and flatters a sustained note far too much. Slice similarity measures 1.000 on a held sine against 0.286 on a plucked phrase: on sustained material every slice is interchangeable, so the stutter has nothing to expose and sounds like a tremolo. Feed it something with transients and pitch variety.
The tape
| control | setting |
|---|---|
span | 400. ms |
heads | 3 |
ratios | 0.25 0.5 1. |
levels | 0.7 0.85 1. |
pans | 0.2 0.8 0.5 |
regen | 0.55 (the ride starts here) |
darken | 5000. |
drive | 0.4 |
wow / flutter | 0.4 0.35 / 0.3 8. |
mix | 35 |
span is defined at the ratio-1.0 head, so the head at 1. returns at
400 ms and the others at 100 and 200. Change span while it runs and the
whole thing glides like tape speed rather than splicing — that is the
transport, and it is the second-best gesture in this rig.
The best one is regen. It goes past 1 on purpose. Past unity the line
self-oscillates, and it stays bounded because the saturator caps what comes
back: the ceiling is |in|max + regen/drive, measured under that value at
every drive tried. So a ride up to 1.4 and back is a controlled build,
not a fire. Keep drive up while you do it — drive 0. removes the
saturator and the cap falls back to unity, which is the setting where a
long ride will not behave.
darken is the generation loss, and it is what makes repeats decay into a
shape instead of just getting quieter. 5 kHz is a good default; below 3 kHz
the tail turns to mud, which is sometimes what you want under a chorus.
The gestures, in order of value
- Ride
regenpast unity and back, whilemixstays put. One hand, whole arrangement. - Automate
densityto 0 and back. Exact bypass, so it reads as the part reassembling rather than a fade. - Move
spanduring a held note. Varispeed glide, not a splice. - Change
seedbetween takes, never during one. reverseandjump. Character, and cheap to overdo. Notejumpis milliseconds, not a probability — at 250 the machine starts quoting material from a quarter-second before the slice it just took, which is where a stutter stops sounding like a stutter and starts sounding like an edit.
When to leave the recipe
- The source is sustained. See the material contract. Put the stutter on the drums and leave the pad alone.
- You want the slices in time with something. Nothing here syncs to a
transport;
stepis milliseconds. Drive it from your own clock if you need bars. - You want the echo to stay clean.
drive 0.gets you a clean line — but then do not rideregenpast 1, because the cap that makes that safe is the saturator you just removed. - You want the dirt to be the point. Then the fuzz belongs last, not
first, and this is a different recipe:
tap.tapecho~→tap.fuzz~with the echo's owndriveat 0. Distorting a wash is a real sound; it is just not this one.
Two hands on a live buffer
tap.scrub~ is the object in this library that most needs a controller
attached before it means anything. Everything below assumes one: an XY pad,
two faders, a trackpad, a phone sending OSC — anything that gives you two
continuous values at once. The recipe is mostly about what to put on each
axis and why.
Measured claims are borrowed from two hands on the same tape.
The chain
source ──┬─▶ tap.scrub~ ──▶ tap.palme~ ──▶ out
└────── (its own dry path, via mix) ──────▶
The scrub records what passes through it, so it goes in the signal path
rather than on a send — there is nothing to send it that it is not already
hearing. Its mix is the dry/wet, and at mix 0 it is the input bit for
bit, which means you can leave it in the chain permanently and have it be
audibly absent until you touch it.
The pad
| control | setting |
|---|---|
maxhistory | 4. s (object argument; bought at DSP start) |
position | X axis, 0–1500 ms |
pitch | Y axis, −12 to +12 st |
drift | 0. |
size | 80. ms |
overlap | 2 |
spray | 0. |
mix | 100 while playing |
smooth | 0. if driving by signal, 20. if by messages |
X is position, Y is pitch, and the whole object exists because those are independent. On tape they would be the same axis — moving the head is the pitch change. Here you can rake back through the last second and a half at the pitch you started at, or hold still and transpose, or do both at once in different directions. Spend the first five minutes doing each separately; the object does not become obvious until you have felt that they do not interact.
size is the texture control. 80 ms is a granular pad. Down at 20 ms
it turns metallic and starts pitching itself at the grain rate; up at
200 ms it stops being granular and becomes a soft varispeed. Sizes that
divide evenly by overlap have an exactly flat window sum — 80 with
overlap 2 does — which matters when you want the still position to be
clean.
overlap 1 is a texture, not a mistake. It leaves gaps between grains:
a gated, chopped version of the same gesture. Worth a switch on the
controller.
The honest bit about pitch
Transposing here warbles, and it is measured rather than apologized for: 98.8 % of a perfect shifter's energy lands within ±15 Hz of the transposed pitch (worst case 91.7 %), so the note is right — what the object loses is concentration, 92.0 % as focused as a clean shift and 75.0 % at worst. Audibly that is a warble, and it is the classic single-delay-line pitch-shifting artifact rather than anything peculiar to this kernel.
Two ways to work with it:
- Lean in.
spray 30.trades the narrow comb for a broadband smear. On sustained material this reads as a texture rather than a fault, and it is the better answer for pads. - Stay out of its way. Keep the Y axis to ±7 and let the position do the work. Small intervals warble least and the object is a scrub pad first.
If you need a clean shift, this is the wrong object — tap.shift~ is built
for it. (tap.pitchaccum~ is built for shimmer rather than transparency,
and has an open issue about
where its line actually sits.)
Freeze, which is the other half
freeze 1 stops the recorder. The playhead keeps going, so the position
now addresses fixed tape and the grains loop the same window — and you can
still scrub, transpose, drift and spray through it. Nothing is going into
the input any more, which is the point: it is a hold you can perform.
A sequence that works on stage:
- Play the phrase through at
mix 0. Nothing happens; the tape fills. mix 100,freeze 1. The last few seconds are now the instrument.- Drag X slowly. This is the scrub.
drift -0.3. The playhead walks backwards on its own while you keep your hand free for Y.spray 40.,size 200.. It stops being a phrase and becomes a pad.freeze 0,mix 0. The room comes back.
Step 6 is exact — mix 0 is bitwise passthrough — so the return is clean
however far out step 5 went.
The one constraint
A grain born position behind the live edge and playing at rate r
reaches position − size·(r−1) behind it by its end. Transpose up with
the position near the live edge and the grain's tail runs off the front of
the tape into the oldest material — a seam.
Nothing clamps it, because clamping would silently bend the pitch to keep
the grain in bounds, which is a worse failure than the seam. Practically:
keep the position at least size·(rate−1) back. At size 80 and one
octave up that is 80 ms. Setting the X axis to start at 100 ms rather than
0 makes the whole problem disappear, and this is why the table above says
0–1500 rather than 0–1500 starting at zero.
What each ingredient buys, in order
- A controller with two continuous axes. Without it this is a delay.
freeze. The half of the object you can build a performance on.size. The texture, and the only control that changes what kind of thing you are playing.- A diffuseur after it.
tap.palme~ @mix 40sustains what the scrub chops; the strings fill the gaps thatoverlap 1opens. spray. Trades one artifact for another. Real, and last.
When to leave the recipe
- You want it in time. Nothing here syncs. That is
tap.stammer~, on the same tape — literally the samecapturecode — and the two are meant to be swapped between rather than combined. - You want a clean delay.
tap.delay~costs a fraction as much and windows nothing. - You want the position to feel like a jog wheel. Put
drifton a spring-loaded control and leavepositionalone: drift is velocity where position is location, and for wheel-like gestures velocity is the right variable.
The instrument in the corner of the room
The Ondes Martenot is not a synthesizer, and the fastest way to make it sound like one is to patch it like one. This recipe assembles the four objects that make up the actual instrument — voice, key, and a loudspeaker with a body — and then spends most of its length on the part that is not a setting at all: what your two hands do.
Everything measured here is borrowed from the instrument that is not a synthesizer and loudspeakers you can play, which in turn cite the circuit paper (Najnudel, Hélie, Roze & Boutin, IEEE/ACM TASLP 28, 2020) and the intensity-key measurement (Quartier et al., Acta Acustica 101(2), 2015).
The chain
[ribbon signal] ──▶ tap.ondes~ ──▶ tap.palme~ ──▶ out
[key signal] ──▶ ▲ (or tap.metallique~)
Two signals in, one instrument out. That is the whole rig, and the temptation to put things between the stages should be resisted until you have played it as it stands — the voice and the diffuseur were designed to be adjacent, and every stage you insert is a stage the real instrument does not have.
The voice
| control | setting | why |
|---|---|---|
ribbon | driven by signal | semitones above A1, not Hz |
key | driven by signal | 0–1 of the physical travel |
depth | 1. | equal oscillators; the full harmonic series |
detect | 0.2 | the published R4 × C21, 200 µs |
drive | 1. | nominal; the harmonics are already there |
keyplacement | 0 | pressure is level |
polarity | 1 | |
power | 0 | the 2A3 moves total harmonics by 0.003 |
oversample | 4 | |
smooth | 0. | the signal inlets are not ramped anyway |
Start there and change exactly one thing at a time, because most of these are citations rather than tastes and the object will tell you when you have left the instrument behind.
The two that are genuinely yours: keyplacement 1 moves the key in front
of the valves, so hard presses get dirty as well as loud — worth about 0.09
of total harmonic content at a half-press, and it is the single change that
most makes the object feel like a synthesizer rather than an ondes.
polarity -1 flips which side of the waveform the preamplifier bends,
worth about 0.12. Try both; keep whichever suits the piece.
depth below 1 is the cheapest real timbre move in the object. At 0.4
the envelope never closes and the tone thins toward a sinusoid — the
closest thing here to a "register", and it is a physical mismatch between
two oscillators rather than an invented control.
The hands
This is the recipe.
The ribbon is linear in semitones, because the circuit paper's Eq. 7
makes it so. That single fact is why an ondes glissando sounds like an
ondes glissando: a hand moving at constant speed produces a constant-rate
glide, not the accelerating swoop a linear-in-Hz control gives you. So
drive ribbon with something that moves linearly in semitones over time:
line~ 0. 36. 4000 ──▶ tap.ondes~ left inlet
Three octaves in four seconds, and it will sound even the whole way. Build
your phrases the same way — line~ or a slow sig~ ramp per note, never a
quantized step unless you specifically want the instrument to sound wrong.
Nothing here rounds to a semitone, and that is deliberate.
The key is the dynamics, and it starts silent. Roughly the bottom 45 % of the travel makes no sound at all. That is the key's own first phase, the elastic strip bending before it reaches the powder bag, and it is why the instrument attacks so sharply: the whole 50 dB lives in the 4.5 mm right after the silence. Practically:
- Drive
keyfrom a pedal, a fader, or aline~— anything continuous. - Expect nothing below about
0.45. If you want the note to speak the instant your controller leaves zero, rescale:scale 0. 1. 0.45 1.. Do that only if you want to give up the attack, because that dead travel is what lets you place an entrance to the millisecond. - The curve steepens through the middle and flattens at the top. Crescendos therefore want a decelerating controller move, not a linear one — which is exactly the feedback a player's finger gets from the real spring.
Play them together. The ribbon without the key is a test tone; the key
without the ribbon is a volume pedal. The instrument is the two hands, and
ondes_ribbon.wav in the render set exists to demonstrate what that
sounds like when both are moving.
The loudspeaker, which is an instrument too
tap.palme~ is the default answer. Twelve strings on a board, and they
sustain what the voice has already stopped playing:
| control | setting |
|---|---|
root | 110. — put the lowest string under your part's key |
tuning | 0 chromatic, so the board answers every note |
decay | 6. |
damping | 4000. |
detune | 4. — cents of scatter, so the board is not a chorus unit |
drive / asymmetry / saturation | 1. / 0.3 / 0.2 |
mix | 60 |
tuning 1 puts the harmonic series on the root instead: the board then
answers only what belongs to that key, which turns a chromatic line into
something that blooms on some notes and stays dry on others. Use it when
the piece really is in one key, and hear it as a compositional decision
rather than a preset.
Watch the level. Twelve resonant loops add up, and a driven board can be a
great deal louder than what went into it — level is there for that, and
it is the one control on these objects you will need to touch first.
tap.metallique~ is the other cabinet: eight plate modes rather than
twelve strings, so it colours instead of harmonizing.
@pitch 180 @decay 6 @tilt 0.8 @brightness 1. @mix 50 is the gong. Push
drive to 3 with asymmetry 0.5 and the distortion happens before the
plate, because that is where the transducer is — a distorted waveform
ringing a gong, not a distorted gong. It is not subtle and it is the most
distinctive sound in the family.
What each ingredient buys, in order
tap.ondes~with both hands moving. Everything else is optional. A static ribbon and a static key is a demo, not an instrument.- The diffuseur. The voice alone is thin on purpose — it is a valve preamplifier output, and it was never meant to be heard without a body after it.
depth. The one timbre control that costs nothing and is physical.keyplacement/polarity. Real, measured, and yours to choose.power. Measured at 0.003 of total harmonic content. Last.
When to leave the recipe
- You want the waveform registers. The real instrument has switchable
timbres — creux, gambe, nasillard and the rest. They are not here, and
they are not here on purpose: no source obtained describes their filter
shapes, and inventing them is the one thing these objects will not do.
If you need those colours, put a filter after
tap.ondes~and call it your filter, not Martenot's. - You want polyphony. The instrument is monophonic and so is this. Two
tap.ondes~in parallel is a duet, not a chord — which is how ondes ensembles actually worked, so it is not a bad answer. - You want the diffuseurs on something else. Take them. A guitar into
tap.palme~is the best argument for shipping them standalone, and nothing about them needs the voice in front. - You want
tap.triode~on its own. It is a published valve stage and it works on anything:@tube 2 @stage 2 @drive 6is the 2A3 power stage used for something it was never in this instrument for.