Introduction

TapTools is a collection of Max/MSP objects with roots back to 1999, rebuilt in 2026 on a portable DSP kernel library (this repository) with thin Max wrappers (the TapTools-Max package). This book is its field guide, in the tradition of the AmbiTap and SampleRateTap books and MuTap's Quieting the Loop: one chapter per object family, written for the person patching at 11 pm, not for the person grading a DSP exam.

Each chapter makes the same promises:

  • It says what the thing is for — and, near the end, when it is the wrong tool, because every tool is sometimes the wrong tool.
  • Every performance claim is measured, not remembered. The numbers in these chapters come from the kernel's own test suite and from the executed verification notebooks in notebooks/, which drive the same C++ code the Max objects compile through a C ABI. When a chapter says "47 dB", a notebook cell measured 47 dB, and you can re-run it.
  • Trade-offs are stated as trades. Knobs that buy something always pay with something; the chapters try to name both sides.

The book is organized the way a patch is:

  • Part I — Sources: the virtual-analog oscillator, tap.vco~, including its analog-character section and the honest Moog recipe.
  • Part II — Filters: the morphing Simper SVF (tap.svf~), the transistor ladder (tap.ladder~), and the envelope filter modeled on the Snow White AutoWah (tap.autowah~) — with its hardware-calibration harness.
  • Part III — Strings, rooms, and spirals: exact true-stereo convolution (tap.convolve~) and the two GRM Tools recreations — the tuned comb bank (tap.5comb~) and the pitch-accumulating shimmer loop (tap.pitchaccum~).
  • Part IV — Tape and time: the Eno recreations — the Discreet Music two-machine tape loop (tap.discreet~), the Music for Airports incommensurate loop bank (tap.airport~), the generative event garden (tap.garden~), and the components they decompose into.
  • Part V — The machines you ride: the Radiohead family — objects whose point is the performance surface rather than a setting. The multi-head tape echo (tap.tapecho~), the live buffer-stutter rig (tap.stammer~), the two-stage fuzz (tap.fuzz~), the granular scrub pad (tap.scrub~), the two Ondes Martenot diffuseurs as standalone driven resonators (tap.metallique~, tap.palme~), and the Ondes Martenot voice itself (tap.ondes~, with tap.triode~ and tap.touche~).
  • Part VI — The spectral set: the 24-band vocoder (tap.vocoder~), the per-bin spectral gate (tap.nr~), and the bin remapper (tap.spectra~).
  • Part VII — The rhythm section: the Roland recreations — the TB-303 voice, its diode-ladder filter, and its sequencer (tap.303~, tap.diode~, tap.303.seq~), and the eight TR-808 voice channels with their row sequencer (tap.808.*, tap.808.seq~).
  • Part VIII — Staying in tune: the pitch corrector (tap.tune~) and the detection/resynthesis machinery it stands on.
  • Part IX — The pedalboard: the stompbox recreations — the voiced feedback overdrive (tap.overdrive~), chasing the TS-lineage feedback pedals rather than a waveshaping curve.
  • Part X — The machine, file by file: the SampleRateTap-style deep dives — one chapter per kernel header, deriving the math, reviewing the code, and recording why each algorithm is written the way it is, alternatives and all. Parts I–IX are for driving the objects; Part X is for trusting them — or changing them.
  • Part XI — Recipes: whole patches chasing specific sounds — the TR-808 kits behind four decades of records, the three-oscillator Moog voice — with settings you can check against the reference pages and the honest accounting of what each ingredient buys.

More chapters land as objects mature; the utility and Jitter objects live in their reference pages, where they belong.

The oscillator and its knobs

Play a perfectly calculated sawtooth for ten seconds and you will learn something uncomfortable: perfection sounds like a diagram. Every cycle identical, every harmonic exactly where the textbook puts it, nothing moving — the ear files it under test tone and stops listening. Now play three of them, each a few cents off the others and each wandering a little, into a filter that pushes back — and the same arithmetic becomes a synthesizer. This chapter is about tap.vco~: what it generates, what each attribute trades, and how to get from the diagram to the instrument — including the honest version of the Moog recipe.

Companion material: the object's reference page (docs/tap.vco~.maxref.xml) and help patcher (help/tap.vco~.maxhelp in the TapTools-Max package) wire up every control in this chapter; the verification notebook shows every number quoted here as an executed, plotted measurement.

One phase, four shapes, no aliasing panic

Inside the object there is a single master phase ramping from 0 to 1 at the frequency you asked for. Everything else is a way of reading that phase: a sine reads it through sin, a saw stretches it to ±1, a pulse compares it to the pulse width, and the triangle integrates the pulse (the classic analog trick, reproduced digitally because it behaves so well). The continuous shape parameter (0 sine → 1 triangle → 2 saw → 3 pulse) crossfades adjacent readings of the same phase, so a shape sweep glides through hybrid waveforms without resetting anything.

The digital oscillator's ancient enemy is aliasing: a naive saw's harmonics march past Nyquist and fold back as inharmonic garbage. tap.vco~ suppresses this with polyBLEP — each waveform discontinuity is rounded across ±1 sample by a polynomial that closely matches what a band-limited step would do. Measured against a naive saw at 3951 Hz (a B7, ugly on purpose): the 13th harmonic folds back to 3.4 kHz, where the naive saw puts it at −27 dB and tap.vco~ puts it at −74 dB — 47 dB of alias suppression right where the ear is most offended.

Two things about this are worth knowing so they don't surprise you:

  • The waveforms look "not band-limited" on a scope. Expected. The BLEP correction touches two samples per edge — at 440 Hz that's 2 of ~109 samples per cycle — and there is none of the Gibbs ripple that brickwall band-limited waves show, because nothing is truncated. The shape stays essentially ideal; the spectrum is what's controlled.
  • Alias suppression is not alias elimination. Push the fundamental into the kilohertz range and distant fold-backs remain, tens of dB down. For melodic and bass registers they are simply gone.

The wiring

   frequency (signal or float)     FM, in Hz (signal)     sync (signal)
              |                          |                     |
        +-----+--------------------------+---------------------+-----+
        |                       tap.vco~                              |
        +-----------------------------+-------------------------------+
                                      |
                              (signal) the waveform
  • Inlet 1 sets the frequency — a float sets the attribute, a signal drives it with true per-sample resolution.
  • Inlet 2 is through-zero linear FM, calibrated in Hz: the input adds directly to the effective frequency. Drive it past the carrier and the phase genuinely runs backward (that's the "through zero" — the classic DX-style sideband sound stays coherent instead of collapsing). Measured: a 500 Hz sine carrier under ±900 Hz of FM stays bounded at exactly 1.0 peak and puts its sidebands where the textbook says.
  • Inlet 3 is hard sync: every rising zero crossing of the input resets the phase, with sub-sample accuracy and an alias correction on the reset. Measured: a 187 Hz slave synced to a 110 Hz master emerges periodic at 110.1 Hz — the pitch follows the master, the timbre follows the slave's frequency, which is the whole trick of sync sweeps.

Single-channel, like every TapTools DSP object: wrap it in mc. for stacks, and keep reading, because the analog section was designed around exactly that.

Signal-flow diagram of the VCO: pitch, FM, and the per-seed analog section sum into a master phase accumulator, which fans out to four waveform readings crossfaded by shape

One phase, many readings — with polyBLEP correcting the edges and the analog section injecting in exactly two places.

The knobs, one by one

frequency, and gliding

Hz, from LFO rates (0.01 Hz) to 20 kHz. Every parameter in the object rides a per-sample ramp whose length is the smooth attribute (ms, default 20) — and on frequency that ramp is portamento. Set smooth to 60–100 ms, send note frequencies as floats, and you have the Minimoog glide, no extra objects. For stepped pitch, set smooth low; for per-sample modulation, use the signal inlet (which bypasses smoothing entirely — you are the smoothing).

shape and waveform

shape is the continuous morph; the waveform sine|triangle|saw|pulse message snaps it to a corner. The corners are the pure shapes; everything between is a crossfade of neighbors on the shared phase. Slow shape sweeps are an underrated modulation destination — the morph is click-free by construction (measured: a 2-second sweep from 0 to 3 keeps its RMS within a factor of ~5 and never drops out).

pw — pulse width

Percent, 1–99, audible as shape approaches 3. The calibration is exact: a bipolar pulse at duty d must average 2d−1, and the measured means at 10/25/50 % are −0.800/−0.500/+0.000. PWM by an LFO into pw (via messages, riding the smooth ramp) is the cheapest "two oscillators" impression one oscillator can give.

gain, presets, interp

gain is output level in dB. Sixteen preset slots store every parameter (store 1 … store 16), and recall morphs to a slot over interp milliseconds (or an explicit time: recall 3 4000) — every parameter riding its ramp simultaneously, shape included. A preset morph across two very different voicings is a patch element in its own right.

seed — which unit you own

Everything random in this oscillator — the drift walk, the jitter noise, and (below) the component tolerances — is generated deterministically from seed. Same seed, same render, bit for bit: your mixes reproduce and the test suite can pin behavior exactly. Different seeds decorrelate. The mental model that pays off: a seed is a serial number. One tap.vco~ with seed 7 is a particular oscillator that came off the line; seed 8 is the unit next to it in the crate. An mc. stack with per-voice seeds is a set of instruments, not copies of one.

The analog section

Here is why a hardware oscillator sounds alive, reduced to what a DSP model can honestly act on. A real VCO is unstable at two time scales — it wanders over seconds (thermal drift) and trembles over milliseconds (noise in the core) — it is mis-calibrated in a structured way (the V/oct converter is exact at its trim point and increasingly wrong away from it), and its waveforms carry the circuit's fingerprints (a bowed ramp, a rounded reset corner, a duty cycle that isn't quite 50 %). None of these is large. All of them are always present, all slightly different from unit to unit, and the ear reads their sum as alive long before it can name any of them.

tap.vco~ models each one with its own control, all in real units, all deterministic per seed, and all exactly zero by default — the default object is the ideal oscillator, and the kernel's test suite pins that at imperfect 0 every seed renders bit-identically.

drift — the slow wander (cents)

A random walk: sample-and-hold noise at ~2 Hz smoothed through a ~0.5 Hz one-pole, scaled to the depth you set. This is the thermal story — the pitch center strolling around over seconds. In a unison stack it is the difference between "chorus effect" and "three players": chorus modulation is periodic and shared; drift is aperiodic and per-voice. Ranges: 3–8 cents reads as a well-serviced vintage instrument; 15–25 as a charming one; 50+ as a broken one.

jitter — the fast tremble (cents)

New with this chapter: the short-time companion — noise at ~80 Hz through a ~40 Hz smoother, so the pitch trembles cycle-to-cycle instead of strolling. Measured at 10 cents depth: the relative spread of individual periods is 2.7×10⁻³ (a few cents, exactly as labeled), against 2×10⁻⁷ for the ideal oscillator — four orders of magnitude more micro-instability, still nothing like vibrato. This is the control that stops a sustained single oscillator from sounding frozen. Ranges: 1–4 cents is felt more than heard; 8–15 is audible grit on pure waveforms.

detune and track — the calibration story (cents, cents/octave)

detune is the static offset — the coarse fact that oscillator 2 was never exactly oscillator 1. track is subtler and very analog: cents of error per octave from A440, the exponential converter drifting from its trim point. Measured with track 5: exactly 0.0 cents at A440, +15.0 cents three octaves up, −15.0 three octaves down, a clean line through the middle. Solo it is nearly invisible; in a stack played across the keyboard it is why vintage unisons get wider — and slightly wilder — up the neck. Ranges: real serviced hardware tracks within 1–3 cents/octave; ±5 is a synth that needs its yearly appointment.

imperfect — the circuit's fingerprints (0..1)

One knob for the waveform-shape story, scaled by per-seed component tolerances so each seed misbehaves in its own direction:

  • the saw ramp bows into the familiar shark-fin (a visible, scope-obvious shape change; spectrally it is mostly a phase effect — stated here so you don't chase magnitude changes that aren't there),
  • the saw's reset corner rounds off: a gentle one-pole closing from ~22 kHz toward ~8 kHz — measured 6.4 dB down at the 40th harmonic (17.6 kHz) at full imperfection; extreme top-end air, traded for warmth,
  • the triangle goes asymmetric, and this one is very audible in the spectrum: the ideal triangle's 2nd harmonic sits at −185 dB (i.e., absent); at imperfect 0.8 it rises to −34 dB relative to the fundamental — even harmonics, the classic "warm" giveaway,
  • the sine picks up mild waveshaper color, and the pulse width takes a small static offset (so two "50 %" pulses from two seeds beat against each other the way two real units do),
  • the whole unit takes a static pitch offset of up to a couple of cents.

Ranges: 0.2–0.4 is a healthy vintage unit; 0.6–0.8 is character you can point to in a mix; 1.0 is a unit with a story. At 0, every seed is the same ideal machine — the analog section never costs you the reference oscillator.

The performance section

Where the analog section models what the circuit does on its own, these controls model what a hand does to it — added after the Recipes chapters had to teach a scaling formula to get constant-width vibrato out of the Hz-calibrated FM inlet.

  • vibrato / vibrato_rate — a sine LFO on the pitch, depth in cents (0–100) and rate in Hz, so ten cents is ten cents in every register. Measured: at a commanded ±100 cents the peak cycle-to-cycle deviation reads 90–110 cents, and the modulation crosses its mean at exactly twice the commanded rate (pinned by test).
  • vibrato_delay — the singing control: the vibrato fades in through a one-pole with this time constant (ms), re-armed on every new note (every frequency-target change), so held notes bloom and passing notes stay plain. Pinned: early deviation under 60 % of settled, and shallow again right after a note change. The signal-rate frequency inlet deliberately does not re-arm — there, you are the modulation.
  • bend — pitch bend in semitones (±24), riding the standard smooth ramp: the wheel, as an attribute. Pinned within 5 cents of the commanded interval.

All of it is deterministic with no randomness — and at depth 0 the output is bit-identical to the ideal oscillator (pinned), so the reference instrument is still free.

The Moog recipe, honestly

The sound everyone wants from this object is three oscillators into a ladder. Here is the recipe, with the honest accounting of which ingredient does what. Rendered A/B demos of exactly this patch (through the real kernels) live in the notebook material.

voicefrequencydetunedriftseed
1f−4 c8 c11
2f+5 c8 c22
3f ÷ 2+2 c10 c33
  • All three: @shape 2 (saw), @smooth 70 for glide, @jitter 3, @track 2, @imperfect 0.3.
  • Sum them (scale by ~1/2.8) and feed tap.ladder~: @mode lp24 @resonance 0.35 @drive 9 @asym 0.45 @comp 0.25.

What each ingredient buys, in order of importance:

  1. The stack itself. Three free-running voices at ±cents is most of the sound. The beating between them is the fatness; the octave-down third voice is the weight. (tap.vco~ free-runs like hardware — no per-note phase reset — so the beat pattern is different on every note. That, not any single voice's tone, is the big analog tell.)
  2. The ladder. drive into the tanh stages compresses and colors the stack; asym adds the even harmonics of mismatched transistors; comp kept low preserves the authentic passband droop as resonance rises. A perfect saw into a driven asymmetric ladder sounds more "Moog" than an imperfect saw into a clean one — spend your character budget here first.
  3. Glide. 60–100 ms of smooth on the note changes. Iconic, and free.
  4. The analog section. Drift keeps the beating from ever repeating; jitter un-freezes sustains; per-seed tolerances make the three voices three units. This is seasoning — essential in the way salt is, invisible in the way salt is.

Omit in reverse order when CPU or taste says so.

When it is not the right tool

  • You need an exact test signal. Actually — it is the right tool: imperfect 0 (the default) is the mathematically ideal oscillator, and the test suite holds it there. Just don't reach for the analog section and a measurement mic in the same patch.
  • You want evolving spectra from one voice — wavetables, granular motion, additive drift. This oscillator's spectrum is fixed per shape by design; morph shape or FM it, but a wavetable oscillator is a different instrument.
  • You want chorus. Twenty cents of drift on one voice is not a chorus; it is a seasick oscillator. Chorus is a delay effect — use one.
  • Noise. The bottom of the shape range is a sine, not a noise source; tap.noise~ has five colors of the real thing.

Checkpoint

One master phase, four shapes and their hybrids, polyBLEP keeping the folded harmonics ~47 dB down. Frequency glides on smooth, FM is in honest Hz and survives through zero, sync locks pitch to the master while timbre stays yours. The analog section is four controls with real units — slow drift, fast jitter, structured mis-calibration in track, circuit fingerprints in imperfect — all scaled by per-seed component tolerances, all exactly off by default, all deterministic: a seed is a serial number. And the Moog recipe is mostly the stack and the ladder — let the oscillator's imperfections season, not carry.

The filter that morphs

Every synthesizer needs one filter it can trust with anything: a bass line, a noise sweep, a parametric EQ move, an audio-rate modulation stunt. tap.svf~ is that filter — a state-variable design in Andy Simper's trapezoidal (zero-delay-feedback) formulation, the same lineage as the filters in Ableton Live, including Auto Filter's Morph type. This chapter is what each attribute trades, and why the design earns the trust.

Companion material: the reference page and help patcher in the TapTools-Max package, and the verification notebook, where every number below is an executed, plotted measurement.

Why "state-variable," and why this one

A state-variable filter computes all its responses — lowpass, bandpass, highpass, notch — from the same two internal states at once, which is what makes continuous morphing between them possible at all. The classic digital version (Chamberlin) famously misbehaves at high cutoffs and under fast modulation. Simper's TPT formulation fixes both: the tuning is prewarped (exact all the way to Nyquist) and the filter is unconditionally stable under per-sample cutoff modulation — the property that later let this same kernel become the sweep engine inside tap.autowah~. The notebook slams the cutoff across five octaves with a 90 Hz LFO under full-band noise; the output stays bounded, no oversampling tricks required.

Signal-flow diagram of the TPT state-variable core: a summing node into two trapezoidal integrators with damping and low feedback, and the output mix that forms every response

Two integrators in a zero-delay loop; every response — and the morph — is three multiplies downstream of the same two states.

The knobs, one by one

type — the discrete responses, the morph, and the EQ family

Ten responses from one core. The classics — lowpass, highpass, bandpass, notch, peak, allpass — plus:

  • morph: one continuous parameter sweeps LP → BP → HP → notch → LP (0 → 0.25 → 0.5 → 0.75 → 1). The corners are bit-identical to the discrete modes — measured max difference exactly 0 — so morphing to a corner is that filter. A slow morph under a held chord is a patch element the discrete modes can't give you.
  • bell, lowshelf, highshelf: the parametric-EQ trio from Simper's coefficient tables, with a ±24 dB gain. Measured: a +12 dB bell peaks at +12.00 dB; a −9 dB low shelf lands −9.00 dB in its plateau and 0.00 dB on the other side. These always run a single 2nd-order section — cascading would square the boost, so order is ignored for them, on purpose.

order — 2, 4, or 8 poles that stay flat

Orders 2/4/8 (12/24/48 dB per octave) run as a cascade with the Butterworth Q spread, so at resonance 0 the response is maximally flat and sits at −3.01 dB at the cutoff regardless of order — measured −3.01 at every one, with slopes of 12.3/24.7/49.4 dB per octave. The trade against a naive cascade of identical sections (which droops long before fc): none. This is just the correct way to stack poles.

resonance — normalized, and honest about the top

0 to 1: 0 is the Butterworth-flat base, 1 is the edge of self-oscillation. Resonance sharpens only the final section of a cascade, so you get one clean resonant peak on a flat passband instead of a compounding stack of peaks. A q message converts to and from engineering Q if you think in those units.

circuit — clean or driven

  • clean is the pure linear filter: cheapest, transparent, never oversampled. Also the reference: the EQ modes and every measured Bode plot above are this circuit.
  • driven adds drive (dB) into a tanh limiter on each section's band node — an OTA-flavored color stage, oversampled (1/2/4×, default 2×). Two measured consequences: a 200 Hz tone through +18 dB of drive grows odd harmonics that simply do not exist in the clean circuit (the 3rd harmonic appears out of the numerical floor, ~140 dB up), and at resonance 1.0 the filter self-oscillates at the cutoff — measured 999.7 Hz for a 1 kHz setting, amplitude bounded by the saturator. It needs a ping to start: a perfectly silent filter is a fixed point.

frequency, the right inlet, and smooth

Float or attribute sets the cutoff through the anti-zipper ramp (smooth, ms). A signal in the right inlet takes over per sample — that's the path for audio-rate filter FM and for envelope-follower patches. Sixteen preset slots morph via store/recall over interp milliseconds, everything gliding together.

Recipes

  • The synth voice: @type lowpass @order 4 @resonance 0.4, envelope into the frequency inlet. Order 4 is the "synth filter" slope; order 2 is the polite one; order 8 is a wall.
  • The DJ sweep: @type morph, sweep morph 0 → 0.5 while easing frequency — the LP-through-BP-to-HP arc is the whole move in one parameter.
  • Tone control: bell/shelves with modest gains. It measures exact, so trust the numbers you type.
  • A sine with character: @circuit driven @resonance 1, ping it, and tune with frequency — a self-oscillating test-tone-with-a-temper.

When it is not the right tool

  • You want the classic squelchy 4-pole growl. That's a transistor-ladder sound — resonance that compresses the passband, saturation inside the loop. Next chapter: tap.ladder~.
  • You need many static EQ bands. One tap.svf~ per band works, but a dedicated multiband EQ (or tap.filter~, the RBJ multimode biquad) is the boring, correct choice.
  • You want the filter to follow your playing. That's tap.autowah~, which is this filter plus an envelope detector and a sweep law.

Checkpoint

One TPT core, every response as an output mix: discrete modes, a morph whose corners are bit-identical to them, and an exact parametric-EQ trio. Butterworth-spread orders stay −3.01 dB flat at any slope; resonance sharpens only the last section; the driven circuit adds tanh color and true bounded self-oscillation at the cutoff. Unconditionally stable under per-sample modulation — which is why other objects build on it.

The transistor ladder

Some filters are tools; this one is a character actor. The four-stage transistor ladder — the Moog circuit — colors everything it touches: the resonance pushes back against the bass, the stages saturate into one another, and at the top of the resonance range it stops filtering and starts singing. tap.ladder~ is a zero-delay-feedback model of that circuit with a tanh saturator in every stage. This chapter is what each control trades, and what the measurements say the model actually delivers.

Companion material: the reference page and help patcher in the TapTools-Max package, and the verification notebook — every number below is an executed measurement. For the linear ladder — the cheap, polite Stilson/Smith model — see tap.fourpole~; this object is its nonlinear sibling.

What the model gets right

Two things separate a serious ladder model from a filter with a "Moog" label:

  • Tuning that survives the top octaves. The classic digital shortcut goes audibly flat as the cutoff rises. This model is prewarped ZDF: measured self-oscillation lands at 1000.2 Hz for a 1 kHz cutoff (0.02 % error) — and, the part that's actually hard, 8009 Hz for an 8 kHz cutoff (0.11 %). You can play the resonance like an oscillator anywhere on the keyboard.
  • Nonlinearity inside the loop, not bolted on. Each stage saturates, and the feedback fights the saturation the way the hardware does. That is where the compression, the "sag," and the bounded self-oscillation come from.

Signal-flow diagram of the ladder: drive into a summing node, four tanh one-pole stages in series, the resonance feedback tap, the comp compensation path, and the Xpander pole-mix taps

The whole filter: four stages, one loop. The red tap sets resonance, the amber paths are the comp bargain and the Xpander mode taps.

The knobs, one by one

frequency and the right inlet

Cutoff in Hz; a signal in the right inlet drives it with true per-sample resolution. Like everything here it rides the smooth ramp when set by message.

resonance — up to and past the edge

0 to 1.1. At 1.0 the loop gain reaches the oscillation threshold; above it the filter sings at the cutoff, amplitude-limited by the tanh stages (ping it to start — silence is a fixed point). Under the edge, resonance does the authentic ladder thing: it eats your passband (see comp).

drive — how hard to lean on the stages

Input gain (dB) into the saturating ladder. Measured THD on a 100 Hz tone: 0.5 % at 0 dB, 3.5 % at 8, 16.5 % at 16, 33 % at 24 — a smooth walk from "slightly thick" to "fuzz pedal's cousin." All odd harmonics, because tanh is symmetric — which is exactly why asym exists.

asym — the even harmonics of real hardware

Real transistors don't match; their operating points sit slightly off-center, and that asymmetry is where a hardware ladder's even-harmonic warmth lives. asym (0..1) models the mismatch. Measured on a driven tone: the 2nd harmonic sits at −156 dB (numerically absent) at asym 0 and rises to −18.6 dB relative to the fundamental at 0.6. One honest warning from the reference page: an asymmetric saturator can produce slight signal-dependent DC — follow with tap.dcblock~ if something downstream cares.

comp — the passband bargain

A real ladder trades passband level for resonance: the feedback subtracts from the input. Measured at resonance 0.9: the passband sits at −13.2 dB with comp 0 (the authentic droop) and at 0.0 dB with comp 1 (fully restored). Vintage behavior or modern behavior — your call, continuously.

mode — pole mixing, the Xpander trick

lp24, lp12, bp12, bp24, hp12, hp24: mixing the ladder's stage taps yields whole families of responses from the same four poles (the Oberheim Xpander's famous trick). Measured small-signal slopes: 23.4 dB/oct for lp24, 11.7 for lp12. The resonance and saturation behavior carries into every mode — a resonant bp24 through drive is a very different animal from tap.svf~'s clean bandpass.

oversample — paying for the saturation honestly

The tanh stages generate harmonics past Nyquist that fold back as inharmonic alias tones. Measured on a hard-driven 5 kHz tone: going from 1× to 4× oversampling drops the non-harmonic (alias) energy by 13.5 dB. The default 2× is the working compromise; use 4× when you drive high notes hard, 1× when you're filtering bass and counting CPU.

solver — fast or exact

The nonlinear loop can be solved with one predictor-corrector pass (fast, the default) or by Newton iteration to convergence (exact, circuit-simulation accuracy). They are audibly identical until drive and resonance are both pushed hard; exact is there for when you want to know, and for renders where CPU is free.

Recipes

  • The bass patch: tap.vco~ saw stack (see the oscillator chapter's Moog recipe) → @mode lp24 @resonance 0.35 @drive 9 @asym 0.45 @comp 0.25. Keep comp low; the droop is the vintage glue.
  • The acid line: @resonance 0.85 @drive 15, envelope into the frequency inlet, and let the resonance fight the saturation.
  • The kick synthesizer: @resonance 1.05, ping it with a click, and ride frequency down fast — a self-oscillating ladder is a sine with attitude.

When it is not the right tool

  • Transparent filtering. Every pole of this filter has an opinion. For surgical work use tap.svf~ (clean circuit) or tap.filter~.
  • Morphing responses. The pole-mix modes switch; they don't glide. Continuous response morphing is tap.svf~'s morph.
  • CPU-constrained patches that just need "4-pole lowpass." tap.fourpole~ is the linear ladder at a fraction of the cost — no saturation, no oversampling, no opinions.

Checkpoint

A prewarped ZDF four-stage ladder with tanh in every stage: self-oscillation in tune within 0.11 % even at 8 kHz, drive that walks THD from 0.5 % to 33 %, asym switching on the even harmonics of mismatched transistors, comp choosing between authentic passband droop and modern flatness, pole-mixed multimode outputs, and oversampling that measurably pays down the saturation's aliasing. The character filter — spend your tone budget here.

The pedal that listens

A wah pedal is a filter with a foot attached. An auto-wah cuts out the foot: it listens to how hard you play and sweeps the filter for you — hit a string and the filter opens; let it ring and the filter settles back down. tap.autowah~ models a specific, beloved instance of the idea: the Mad Professor Snow White AutoWah, Björn Juhl's OTA-based envelope filter, grounded in the traced circuit and the published behavior. This chapter is how to drive it — and how we will know, measurably, when the model matches the pedal.

Companion material: the reference page and help patcher in the TapTools-Max package; the design document (plans/tap.autowah~.md in TapTools-Max) with the full hardware research; and the validation notebook, which measures everything below and ends with a cell waiting for recordings of the real pedal.

What the hardware is, in one paragraph

A 2-pole state-variable filter (an LM13700 OTA circuit) whose frequency is pushed up from a resting point by an envelope detector — a diode and a capacitor, charged fast, discharged at a rate you set. Four knobs: Sensitivity (how hard your signal drives the sweep), Decay (how fast it falls back), Bias (the resting frequency), Resonance (the Q). Published sweep: 250 Hz to about 2.5 kHz — a throaty, vocal range, deliberately unlike the quack of a Mu-Tron-style filter. One secret feature: with Sensitivity at minimum it becomes a fixed, manually swept filter — the "cocked wah."

The model composes tap.svf~'s Simper core (one 2nd-order section, driven per sample — the modulation stability that filter chapter promised, cashed in) behind a rectifier → attack/release follower and an exponential sweep law. The measured control behavior:

  • Sweep law: cutoff = bias · 2^(sweep · range). Measured against the design curve across the full envelope range: max error 0.000 cents. The law lives in one function on purpose — if the real pedal turns out to sweep linearly in Hz, one function changes and nothing else moves.
  • Timing: attack set to 2 ms measures 1.94 ms; decay set to 250 ms measures 256 ms, and the release fits a pure exponential with residual σ = 0.004 — an RC discharge, like the hardware.

Signal-flow diagram of the autowah: the audio path through the borrowed SVF core and dry/wet mix, with the detector chain of sensitivity gain, rectifier, RC follower, tanh knee, and exponential sweep law driving the cutoff

A detector, a law, and a borrowed filter — the amber chain is everything this object adds to the SVF it composes.

The knobs, one by one

sensitivity — the trigger level, and the secret mode

Detector input gain in dB (−60..+24). Tune it to your instrument and touch: too low and only your hardest hits open the filter; too high and everything pins at the ceiling (a tanh soft knee compresses hard playing into the top rather than slamming a rail). At −60 the envelope is exactly off and the object becomes the cocked wah: a fixed resonant filter with bias as the manual sweep control. Factory preset 4 ships that voicing.

decay — the personality knob

How fast the filter falls back to bias, in ms (10..5000). Fast (tens of ms) gives a wah articulation on every note — the funk setting. Slow (hundreds of ms up) gives classic auto-wah swells that ride your phrasing. This is the knob to perform.

bias and range — where the sweep lives

bias is the resting frequency (default 250 Hz, the hardware's home); range is the sweep span in octaves above it (default 3.3, the hardware's 250 → ~2500 Hz). Both go far beyond the hardware if you want them to, and direction 1 sweeps down from bias instead — a TapTools extension the pedal never had.

resonance, mode, drive

resonance (0..1) is the Q — the vocalness. mode picks the filter tap: lowpass is the stock voicing; bandpass is the circuit's other node, a known hardware mod — quackier and noticeably quieter. drive (dB) engages the saturating SVF circuit for OTA-flavored color; 0 keeps it pure.

attack, mix, and the rectifier

The hardware's attack is fast and fixed; ours defaults to the same 2 ms but is exposed (0.05..100 ms) for softer onsets. mix is an equal-power dry/wet the pedal never had — 100 % (wet-only) is the hardware. And under the hood the detector's rectifier is selectable: full-wave (default, cleaner tracking) or half-wave (the traced single-diode topology). Measured: the half-wave detector carries 2.6 % signal-rate ripple on a low tone against the full-wave's 0.7 % — a real, quantified flavor difference awaiting the hardware A/B.

The sidechain inlet and the envelope outlet

A signal in the right inlet takes over the detector: one sound wahs another (a kick opening a pad is the classic). The right outlet emits the envelope (0..1) as a signal — the detector as a free modulation source for anything else in the patch.

Factory voicings

Preset slots 1–4 ship guitar, bass (lower bias, tighter range — the GB pedal's instrument switch, as a preset you can morph to), slow swell, and cocked wah. recall 2 4000 morphing from guitar to bass over four seconds is its own effect.

How we'll know it matches

The validation harness is built and proven on ground truth. An STFT peak-trajectory extractor recovers the swept resonant peak from wet audio alone — no dry reference needed — and against the kernel's own cutoff trace it correlates at 0.979 in log-frequency, with the small measured offset (−37 cents) close to what resonant-peak physics predicts (−20 cents at that Q). The same code runs on demo videos and on the real pedal. A Snow White is on order; when it arrives, reamped recordings drop into notebooks/reference/ and the notebook's last cell overlays hardware against model. Disagreements map one-to-one onto kernel constants. That pass may flip the default filter tap or the sweep law — both are flagged, isolated, and waiting.

When it is not the right tool

  • You want the filter on a knob or LFO, not your dynamics. That's tap.svf~ with a signal in its frequency inlet — this object's own core, without the detector.
  • Your source has no dynamics. A static pad through an auto-wah is a static filter. Feed the sidechain something rhythmic instead.
  • You want the Mu-Tron quack. Different circuit, different voicing — raise resonance, try mode 1, but know you're modding a Snow White, not summoning a Mu-Tron.

Checkpoint

An envelope detector with a fast attack and a musician's decay knob, driving a 2-pole resonant filter up from bias through an exponential law that measures exact to the design. Sensitivity at the floor is the cocked wah; the sidechain inlet and envelope outlet make the detector patchable; the factory slots hold the four voicings that matter. And the model doesn't ask to be trusted — the extractor that will judge it against the real pedal is already built, already proven, and already waiting in the notebook.

Borrowed rooms

Every room you have ever heard is a filter: clap once, and what comes back — the impulse response — is everything the room will ever do to any sound. Convolution reverb plays your signal through that recording. tap.convolve~ does it in true stereo against an impulse response held in a buffer~, using the standard engine of the genre (uniformly-partitioned overlap-save FFT convolution), and its defining property is worth stating up front: it is exact. Not "high quality" — exact. This chapter is what that buys, what it costs, and how to drive the two knobs that actually matter.

Companion material: the reference page and help patcher in the TapTools-Max package, and the verification notebook — the first notebook in this repo, and the template for all the others.

Exact, measured

The engine splits the IR into blocksize-sample partitions, transforms each once, and multiply-accumulates in the frequency domain over a delay line of past input spectra. The bookkeeping is intricate; the result is not. Measured against a direct time-domain convolution of the same float32 IR: maximum difference 3×10⁻¹² — double-precision noise. Change the block size and the output doesn't change either (64/256/1024 agree within 2×10⁻¹², latency-removed). An impulse through the engine reconstructs the IR to 5.5×10⁻¹⁴, and a synthetic 0.60 s-RT60 reverb measures back at 0.599 s. There is no "character" in this engine to audition; the character is entirely in the IR you load.

The wiring, and the one real cost

Stereo in, stereo out, IR from a named buffer~. The cost: latency of exactly blocksize samples — verified for every block size — on top of your I/O latency. That is the entire quality/latency dial:

  • blocksize small (64–128): tight enough for live input; more CPU per sample (more, smaller FFT batches).
  • blocksize large (512–2048): cheapest; latency grows to match. For a send/return reverb on a mix bus, nobody hears 21 ms of pre-delay you didn't ask for — except you, so use predelay deliberately instead.

maxsize reserves capacity (both are locked in while DSP runs and apply on restart). The rest of the surface mirrors tap.verb~ so the two reverbs read as siblings: mix, gain, predelay, normalize (energy-based, so quiet and hot IRs land at comparable levels), bypass, mute.

Diagram of uniformly partitioned overlap-save convolution: framed input blocks through an FFT into the frequency-domain delay line, multiplied per bin against the double-buffered IR partition spectra, then IFFT with the aliased half discarded

The wiring: one FFT in, one IFFT out, and a multiply-accumulate that is the only cost growing with IR length.

True stereo, by channel count

A stereo room isn't two mono rooms: sound from the left source arrives at the right ear too. The engine runs the full 2×2 matrix — LL, LR, RL, RR — and the buffer~'s channel count selects the topology:

  • 4+ channels: true stereo, all four paths (measured: a signal sent only left emerges on the right at exactly the cross-feed path's gain, 0.600 expected, 0.600 measured, with zero leakage where paths are silent).
  • 2 channels: dual mono — L and R convolved separately, no cross-feed.
  • 1 channel: the same mono room on both sides.

Loading rooms while the music plays

IRs are analysed off the audio thread and published atomically into a double-buffered slot: swapping IRs mid-performance neither clicks nor drops (measured RMS across the swap instant: 21.9 before, 22.1 just after), and one block later the output is bit-identical to an engine that had the new IR from the start. Load rooms like presets; the engine doesn't flinch.

Recipes

  • The honest room: a measured IR (church, plate, spring — the internet is full of them), @mix 25 @normalize 1, and resist the urge to EQ the IR itself before trying predelay — 10–30 ms of it buys clarity for free.
  • Not a reverb at all: an IR is any filter. A single click is a delay; a strummed guitar body is a body simulator; a vowel is a formant filter. Convolution doesn't know it's supposed to make reverb.
  • True-stereo width: record or synthesize the four paths with a genuinely different LR/RL from LL/RR — the cross-feed is where "being in the room" lives.

When it is not the right tool

  • You want to design the reverb — decay knobs, damping, modulation, gated tails. A static IR can't do any of that; tap.verb~ (the algorithmic Moorer reverb, with its own oversampling and limiter) can.
  • Zero-latency insert on a live path. The engine costs blocksize samples, full stop. At 64 that's 1.3 ms — small, not zero.
  • Time-varying convolution. Swaps are click-free but discrete; the engine doesn't interpolate between rooms.

Checkpoint

Partitioned convolution is exact linear convolution — measured to 10⁻¹² — with one honest cost, blocksize samples of latency, and one honest dial, block size against CPU. True stereo comes from the buffer's channel count; IR swaps are atomic and dropout-free; and everything the effect sounds like is the impulse response you feed it. Borrow better rooms.

Five strings, no guitar

Feed a comb filter its own output and it stops being an EQ curiosity and becomes a string: a resonator with a pitch, a ring time, and a temperament. Five of them, tuned by hand, is an instrument — that was the insight of the GRM Tools Classic "Comb Filters" plugin, and tap.5comb~ is its recreation: five resonant combs with per-voice tuning, masters that play the whole bank, and the preset-morph engine that made the GRM tools feel alive. This chapter is how to tune, ring, and morph it.

Companion material: the reference page and help patcher in the TapTools-Max package, and the grm_comb_render tool in the kernel repo, which renders the listening-check scenarios outside Max.

What a resonant comb actually is

A delay of 1/f seconds fed back on itself resonates at f and all its harmonics — pluck it with noise and it rings like a string tuned to f. Two implementation details decide whether five of them sound like a chorus of strings or like a broken flanger, and both were the reasons this object was recreated rather than ported:

  • Fractional delays. At 48 kHz, a 440 Hz comb needs a delay of 109.09 samples. Round it to 109 and the comb plays 440.37 Hz — every voice lands on a slightly wrong, slightly different wrong pitch, and the beating between voices (the whole point of a bank) is gone. The delays here are Hermite-interpolated: the tuning is continuous, and sweeps glide instead of zippering.
  • No clipper in the loop. The feedback path uses a DC blocker and a precise feedback cap, not a hard limiter — high resonance rings clean instead of distorting.

Signal-flow diagram of one comb voice: input sums with feedback into a fractional delay line, read by Hermite interpolation, with the feedback ring running through the loop lowpass, normalized DC blocker, warp allpass, and ring-time-derived feedback gain; a pickup tap at half the loop feeds the output subtractor

One voice of five. The red ring is the string; the amber tap is the pluck position.

The knobs, one by one

freq1..5 and freq — the tuning

Per-voice frequencies (5 Hz floor, the GRM's own) — or the notes message, which tunes up to five combs from MIDI note numbers in one list (fractional allowed) — plus a master multiplier (0..2) that transposes the whole bank — the master is the performance control, gliding every voice proportionally so chords stay chords.

res1..5 and res — ring time, not feedback

Resonance is mapped to ring time on a log curve, 20 ms to 100 s, and the feedback coefficient is derived from the current delay — so a voice keeps its ring time as its pitch sweeps, instead of ringing longer at low notes and choking at high ones (the raw-feedback behavior of naive combs, and of the legacy abstraction). 50 is a decaying pluck; 80+ sustains; near 100 it is a drone that outlives your patience.

lp1..5 and lp — the string's brightness

A one-pole lowpass inside each feedback loop: every pass around the loop gets darker, which is exactly how real strings decay (highs first). Open it for metallic; close it toward a few hundred Hz for felt and thump.

warp — stiff strings

New to this recreation: a negative-coefficient allpass in the loop disperses the partials — upper harmonics round-trip faster and stretch sharp, the inharmonicity of a stiff piano string. The main tap is compensated at each voice's fundamental, so the pitch stays put while the timbre goes piano-ish, then bell-ish. At extreme warp × high tuning the loop can't get shorter than the dispersion and the pitch flattens — physical, and documented.

phase — where you pluck the string

Also new: a half-loop pickup tap. At 100 the even harmonics cancel — the sound of plucking a string exactly at its midpoint. Neutral at 0.

gain, mix, and the morph engine

Equal-power dry/wet and output gain, plus the GRM hallmark: sixteen preset slots with timed interpolation. store 1, retune everything, store 2, then recall 1 8000 — every frequency, resonance, and damping glides for eight seconds through territory you never explicitly tuned. Grabbing one parameter mid-morph overrides just that parameter. The morph is not a transition between sounds; it is the sound.

Recipes

  • The resonator chord: tune freq1..5 to a voicing (say 80/120/160/200/ 102 Hz — the legacy factory preset), res high, and feed it drums or speech. The input is now an excitation signal for your chord.
  • The piano that isn't: moderate resonance, warp 40, lp around 3 kHz, and pluck with clicks.
  • The eight-second gesture: two stored extremes and a long recall — the classic GRM move. Automate nothing else.

When it is not the right tool

  • One comb, precise and plain: tap.comb~ is the single, cheaper unit.
  • Echoes rather than pitch: delays long enough to hear as repeats are tap.delay~ / tap.multitap~ territory — a comb is a delay, but this one is tuned and normalized for resonance, not slapback.
  • Faithful nostalgia: this deliberately is not the legacy tap.5comb~ abstraction — the integer delays, linear feedback, and in-loop clipper it had are exactly what was retired, and the deviations are flagged in the reference page for the audition.

Checkpoint

Five Hermite-tuned resonant combs with ring time on a log map (20 ms–100 s), per-loop damping, and two ways to bend the string physics (warp for stiffness, phase for pluck position) — under masters that transpose the bank and a sixteen-slot morph engine that turns retuning into performance. Strings, chords, drones, and gestures; no guitar required.

The spiral staircase

Most pitch shifters are a one-way trip: in, transposed, out. The GRM Tools "PitchAccum" closed the loop — the transposed signal is delayed and fed back into the transposer, so every pass around the loop shifts it again. Set +7 semitones and a note becomes a rising spiral: +7, then +14, then +21, each echo climbing, the whole thing dissolving upward like light on water. That loop is the effect everyone now calls shimmer, years before the name. tap.pitchaccum~ is the recreation: two independent transposer-delay loops ("shadows") with the accumulation wired in. This chapter is how to climb.

Companion material: the reference page and help patcher in the TapTools-Max package, and the grm_pitchaccum_render tool in the kernel repo for listening checks outside Max.

The loop, and why it doesn't collapse

Each shadow is: granular transposer (±24 semitones) → delay (up to 3 s) → feedback → back into the transposer. Three design choices keep the spiral musical instead of muddy:

  • Constant-level grains. The transposer sweeps two taps half a cycle apart, each windowed so the pair sums exactly to 1 at every phase and every crossfade width. The original tt_shift engine's window pair didn't quite sum flat, which imposed an amplitude ripple at the grain rate — after ten trips around a feedback loop, ripple compounds into tremolo. Here the tenth pass is as steady as the first.
  • Hermite-interpolated taps. Fractional delays keep each pass in tune, so the spiral's steps are the interval you set, not the interval plus drift.
  • A capped, DC-blocked loop. Feedback tops out at 0.99 with a DC blocker in the path — the spiral can run for a very long time, but it is unconditionally bounded (unit-tested at the cap).

The signature is measurable: set +7 semitones and the kernel test finds energy at +7 and +14 — the second pass, the accumulation itself.

Two signal-flow diagrams contrasted: tap.pitchaccum~ feeds its output back into the delay buffer upstream of the moving transposer taps, so every pass is transposed again; the ordinary shifter-in-a-feedback-loop patch taps the delay output and only ever shifts once

The topology is the effect. Feedback re-enters upstream of the taps, so the staircase climbs; in the ordinary patch every echo is the same interval.

The knobs, one by one (per shadow, ×2)

trans1 / trans2 — the step of the staircase

±24 semitones, continuous. Musical intervals (+7, +12, +5) make harmony; small offsets (±0.1–0.3 st) make lush detune-echo instead of a spiral; negative values descend into the dark version nobody expects.

delay1 / delay2 — the tread depth

Up to 3 s per shadow. Short (50–150 ms) blurs the passes into a texture; long (0.5–2 s) articulates each step of the climb as an audible echo.

fb1 / fb2 — how many steps

How much survives each trip, 0–99. 30 gives two or three audible generations; 70 a long climb; 90+ a texture that essentially sustains until the transposition walks it out of range (energy shifted past the audible band is the spiral's natural exit).

xfade — the grain crossfade

GRM's Cross-fade control: the width of the grain envelope's flanks. Narrow is more articulate and more grain-rate flavored; wide is smoother and softer in attack. Because the envelope pair always sums to 1, this changes texture, never level.

The modulation section, and follow

A global LFO (with modphase offsetting shadow 2, so the two loops breathe against each other) plus per-voice deterministic random transposition modulation — a little of either keeps long spirals from sounding cloned. follow (off by default) engages a pitch follower — decimated normalized autocorrelation, confidence-gated, deliberately picking the smallest plausible lag so it doesn't lock onto subharmonics — which adapts the grain window toward the detected period: cleaner transposition on monophonic sources, ignored gracefully on noise.

Presets

The sixteen-slot morph engine, as everywhere in the GRM pair: two stored spirals and a timed recall between them is a gesture in itself.

Recipes

  • Shimmer, the classic: shadow 1 at +12, delay ~400 ms, fb1 75; shadow 2 at +7, delay ~650 ms, fb2 60; both into a reverb (tap.convolve~ with a long church, or tap.verb~). The reverb is load-bearing — shimmer is spiral plus wash. The full patch has its own recipe in Part IX.
  • The descent: −5 and −12, long delays, moderate feedback — a staircase into the basement, much rarer and much creepier.
  • Micro-thickener: ±0.15 st, 60/90 ms delays, feedback 50, xfade wide — not a spiral at all, just an expensive-sounding widener.

When it is not the right tool

  • One clean transposition, no loop: tap.shift~ is the plain shifter — same modernized engine, none of the plumbing.
  • Formant-true vocal shifting: granular transposition shifts formants with the pitch; chipmunks live this way. tap.harmony~ is the dedicated tool — formant-preserving voices at fixed intervals, chords included.
  • Rhythmically exact multi-tap echoes: the delays here serve the loop; tap.multitap~ serves the grid.

Checkpoint

Two transposer-delay loops where the feedback re-enters the transposer, so pitch accumulates pass after pass — +7 becomes +14 becomes +21. Constant-sum grain envelopes keep the tenth pass as steady as the first; Hermite taps keep it in tune; the capped, DC-blocked loop keeps it bounded forever. Intervals are the architecture, delay is the pacing, feedback is the height — and the morph engine turns the whole staircase into something you can bend mid-climb.

The tape that forgets slowly

Every other delay in this house is kept honest by a cap: feedback stops just short of one, because a loop that gains nothing and loses nothing will pile up until it clips. tap.discreet~ is built on the opposite bargain. Its regeneration goes all the way to 1.0 — legally, cleanly, forever — because the loop forgets: every pass through the tape comes back a little darker and a little softer than it went in. The memory loss is not a defect the kernel tolerates; it is the mechanism that keeps the machine stable. You are not patching a delay effect. You are renting a machine whose memory is the instrument.

The rig it recreates is printed on the back cover of Discreet Music (Obscure/EG, 1975): Brian Eno's synthesizer feeding one Revox tape machine, the tape spooling for seconds across the room to a second machine, and the second machine's playback both sent to the speakers and folded back into the first machine's record head. It is the same two-machine system Robert Fripp ran for the No Pussyfooting loops. The tape path itself — the fractional read, the periodic wow and flutter, the in-loop coloration — follows the published tape-echo modeling literature (Arnardóttir, Abel, and Smith's AES model of the Echoplex, and Välimäki et al.'s tape-echo work). The schematic is the score; this kernel is a faithful performance of it.

Companion material: the executed notebook discreet.ipynb, which measured every claim below, and the eno_render tool, whose discreet_basic and discreet_sustain scenarios are the listening copies. The Max wrapper lands in the TapTools-Max package alongside the rest of the family.

Signal-flow diagram of tap.discreet~: input through a send-level fader and record head onto seconds of tape, a wow/flutter-modulated play head, an equal-power dry/wet mix out, and a red return path of darkening lowpass, bounded saturation, DC blocker, and regeneration gain back into the record head

Two machines and a spool of tape; the red return is where the forgetting — and therefore the stability — lives.

loop — the tape span

loop_seconds is the distance between the machines: how long a phrase travels before it returns. The kernel test pins the grid to the sample — an impulse comes back at exactly one loop, bit-for-bit the first time, and every later return lands within a sample of its grid point.

Changing the loop while audio runs is a tape-speed change, not a menu option: the read head physically glides to its new distance, and gliding a read head is doppler. Move from 0.5 s to 0.75 s over half a second and the playback drops an octave while the transport re-spools, then re-locks on pitch — the test measures 220 Hz mid-glide and 440 Hz within five cents after. There is no crossfading "digital" mode, on purpose. If a pitch bend on loop changes would ruin the patch, this is the wrong delay (see below).

regen is the return level into the record head, and unlike tap.delay~'s feedback (capped at 0.99), it reaches exactly 1.0. The notebook plays a one-second noise burst into the loop at regen 1.0 and lets it run for twenty seconds: the level settles and stays — no growth, no collapse — because the wear path bounds it. The saturator's output can never exceed 1/drive regardless of what the loop accumulates, the DC blocker keeps offsets from stacking, and the darkening lowpass decides what survives: lows sustain, highs surrender. The pinned scenario is blunt about the contract — it asserts non-growth, never decay, because at regen 1.0 sustain is the promise. Bring regen down, or darken harder, to end a piece; clear is the eject button, and regen-1.0 material is gone for good.

darken and drive — the wear

darken_hz is the record/playback corner: every pass through the loop runs through a one-pole lowpass at this frequency, so a bright phrase sheds its treble generation by generation while its body lingers. This is measured, not vibes: with the corner at 2 kHz, a 6 kHz tone loses to 0.292 of itself per pass and a 300 Hz tone keeps 0.890 — and both numbers match the analytic transfer of the wear path to three decimals in the executed notebook.

Per-pass level of a 300 Hz and a 6 kHz tone recirculating through the loop, measured points landing on the analytic prediction lines

Generation loss, measured against regen · |H_wear|. The tape forgets treble first.

drive is the record-head saturation — the guarantee. At any drive above zero the loop is absolutely bounded no matter the settings; at drive 0 the path is exactly linear (a real bit-for-bit passthrough, not "almost") and the loop leans on darkening alone. Drive around 0.5 is the tape sound; drive high is the loop slowly compressing itself into a wash.

wow and flutter — the transport

Two sines, slow-deep and fast-shallow, breathing the play head's position. The pitch math is honest and checkable: depth times 2π times rate is the peak deviation, so 2 ms of wow at 0.5 Hz predicts ±10.9 cents — and the notebook's YIN pitch track measures 10.9. The transport is periodic and deterministic by design (no stochastic capstan drift): two renders of the same settings are bit-identical, which is also a pinned test. Set both depths to 0 for a perfectly still machine.

input_level — the performance move

The fader Eno actually rode was not the output — it was the send. Play a few phrases into the machine, then bring input_level to zero: the loop keeps unrolling everything it holds, worn a shade further every pass, and the piece continues without you. That gesture — set up a system, feed it, step away — is the whole record, and it is one setter here. mix is the ordinary equal-power dry/wet with bitwise-exact endpoints.

Recipes

  • The Discreet Music bed: @loop 5. @regen 0.95 @darken 3500 @drive 0.4 @mix 60. Play sparse, slow phrases; stop; listen to what the tape decides to keep.
  • Frippertronics: @loop 6.5 @regen 1. @drive 0.7 @darken 2200 @mix 100. Solo over yourself from a minute ago. The wash never clips and never ends until you end it.
  • Haunted slapback: @loop 0.15 @regen 0.85 @wow 4. 0.9 @flutter 0.15 12. — a short loop with a seasick transport; the doppler and the wear turn a slap delay into a memory of one.
  • The exit: whatever is running, ride @regen from 1. to 0.7 over a minute. The piece performs its own fade, oldest material first.

When it is not the right tool

  • Rhythmic delays. Loop changes bend pitch by design, and there is no tempo sync. tap.delay~ is the clean line; tap.multitap~ is the pattern.
  • Anything that must not color the repeats. Wear is always in the loop (drive 0 removes only the saturation, not the darkening you set). If the tenth echo must equal the first, this machine is philosophically opposed.
  • Loops that should line up with other loops. One machine, one spool. For a bank of independent free-running loops, the next chapter's tap.airport~ is the instrument.

Checkpoint

Seconds of tape between two machines; a worn return path — darken, saturate, DC-block — instead of a feedback cap; regeneration to exactly 1.0 because forgetting is the stabilizer. Loop moves are honest tape-speed doppler, the transport is two deterministic sines measured in cents, and the send fader is the performance. Every number above lives twice: as an executed cell in discreet.ipynb and as a pinned scenario in tests/discreet_test.cpp, which CI runs on every push.

Loops that never line up

Take seven tape loops of deliberately awkward lengths — none a multiple of another — put one soft phrase on each, and let them all turn at once. Each loop is trivial: it plays the same thing forever. The system is not: the phrases drift against each other, meet, part, and meet again differently, and the pattern of coincidences does not repeat within a human afternoon. That is "2/1" from Brian Eno's Music for Airports (Ambient 1, EG, 1978), as he described the rig in the album's liner notes and in A Year with Swollen Appendices: the lengths are the score, and the machine's whole job is to keep the loops turning without an opinion. tap.airport~ is that machine — up to eight free-running loops, each with a single head that both plays and records, summed to stereo.

The discipline that makes it the instrument it is: nothing resets a phase. Not recording, not a level move, not a pan, not even a length change. The free-run is the composition, and the kernel treats the heads as sacred; the test suite literally hammers every setter mid-run and then checks that the heads have advanced by exactly the samples processed.

Companion material: the executed notebook airport.ipynb, which measured every claim below, and the eno_render tool's airport_two_one scenario — three stereo minutes of seven loops, the listening copy. The Max wrapper lands in the TapTools-Max package alongside the rest of the family.

Signal-flow diagram of tap.airport~: one tape loop of eight drawn as a circle with a single play-and-record head, input through a record gate, playback through darken, level, and equal-power pan into stereo sums, with the other loops ghosted behind

One loop of eight. The head plays, then records, then advances; nobody ever tells it where to be.

Record and return

record(loop, 1) punches the input onto that loop's tape at wherever its head happens to be — there is no downbeat, no quantized punch-in, because Eno's rig had none. Recording replaces (each phrase was recorded once, not overdubbed), and playback reads just ahead of the write, so while recording you hear the previous generation under the head. record(loop, 0) freezes the tape, and freezes it bit-exactly: the pinned test compares two whole passes of a frozen loop and requires them identical to the bit. A loop is not a degrading medium here — it replays the same magnetic imprint every revolution, which is why this kernel deliberately has no per-pass generation loss (that is tap.discreet~'s physics, not a loop's).

The lengths are the score

length_seconds per loop is where the composing happens. Two loops of 24000 and 30000 samples realign only at their least common multiple — 120000 samples, 2.5 seconds — and the kernel will tell you: composite_period_seconds reports exactly 2.5 for that pair, confirmed in the notebook by rendering the coincidence raster and watching it repeat at 2.5 s and at no shorter lag.

Return raster of two incommensurate loops and their sum, with the 2.5-second composite period marked

Two awkward lengths and their coincidences. Stretch the lengths and the composite period leaves the room.

Then stretch toward the piece: give seven loops airport-scale lengths in awkward ratios and the composite period overflows a 64-bit sample count — the kernel reports infinity, which is not a failure mode. It is the point.

Changing a length while running is a splice: the tape keeps its content and the head re-wraps modulo the new length — never rewinding — exactly as cutting a physical loop shorter would land you mid-phrase. It can click. Splices do.

Level, pan, shade

Each loop has a slewed linear level, an equal-power pan with exact endpoints (a hard-panned loop is bitwise absent from the far bus — the same law as tap.multitap~), and a darken corner that shades that loop's playback tone. The shade is a static one-pole per loop, not wear: measured in the notebook, a 6 kHz phrase through a 1 kHz shade lands at 0.169 of its transparent twin, against an analytic prediction of 0.169. At the band ceiling — the default — the shade stage is bypassed entirely and playback is bit-transparent, which is what makes the freeze and hard-pan promises testable as bitwise facts rather than tolerances.

There is deliberately no wow here: the phasing engine of "2/1" is the incommensurate lengths, not pitch drift. If a loop's source should breathe like tape, run it through tap.discreet~ on the way in.

Recipes

  • The terminal: seven loops, @lengths 17.8 19.1 21.3 23.9 26.2 28.7 30.9, one sustained tone phrase recorded onto each, levels around 0.45, pans spread wide, a 4 kHz shade on two of them. Let it run. Come back in an hour; it will not have repeated.
  • Phase study: two loops, lengths in a near ratio (say 8.0 and 8.1), the same short phrase on both, panned hard left and right — the Reich-adjacent version, where the drift itself is the melody.
  • Sound-on-sound sketchpad: one loop, @lengths 12., record gate on a footswitch. Punch in fragments as they occur to you; the head's indifference to your downbeat is the charm.
  • Breathing loops: patch sources through tap.discreet~ (gentle wow, regen 0) before the record gate — tape transport on the way in, stable free-run once captured.

The same machine, in pieces

There was never a loop bank doing loop-bank things in here — there is an array of eight identical lanes and a summing loop. That lane is now an object of its own, tap.reel~, and three of them summed are a tap.airport~ bitwise (pinned in tests/airport_test.cpp). Patch it instead of using this object when you want an insert on one loop, a varispeed on one reel, more than eight loops, or tape you actually use — the bank buys all eight worst-case reels at DSP start regardless. See The same machine, in pieces.

When it is not the right tool

  • Synchronized looping. This machine never lines up by design. A beat-locked looper wants a phase reset on the downbeat, which is the one thing this kernel refuses to do.
  • Degrading loops. A frozen loop here is bit-eternal. For material that should wear out as it circulates, tap.discreet~ is the machine with the forgetting built in.
  • Dense delay textures. Eight long loops is a composition system, not an echo; tap.multitap~ does a hundred taps without ceremony.

Checkpoint

Up to eight free-running loops, one sacred head each: record replaces at wherever the head is, freeze is bitwise, splices re-wrap and never rewind, and no setter touches a phase. Level, exact-endpoint pan, and a bypassable playback shade place the phrases; the lengths do the composing, and composite_period_seconds tells you how long until the piece repeats — ideally, longer than you will be alive. Every number above lives twice: as an executed cell in airport.ipynb and as a pinned scenario in tests/airport_test.cpp, which CI runs on every push.

The garden that plays itself

The first two chapters of this part recirculate sound: tape that forgets, loops that never agree. This one recirculates decisions. Plant a note and it comes back every pass of the loop a step quieter and a step purer, until it fades below hearing and retires. Plant several and they braid. Stop planting altogether and, after a patient interval, the garden starts planting for itself — always on the scale, never in a hurry. You do not play this instrument so much as tend it, which is exactly the posture Eno kept asking for: the composer as gardener, not architect. The kernel is named for that metaphor.

What it recreates is the principle behind Brian Eno and Peter Chilvers' generative apps (Bloom, 2008), as described in their published interviews and in Eno's 1996 "Generative Music" talk: touch becomes note, note repeats and fades, scale makes wrong notes impossible, idleness hands the piece to the system. The principle only — no scale tables, timings, or sounds are taken from the app, and its name is a live trademark of Opal Limited, which is why this object is a garden and not a bloom. (As with tap.tune~'s history paragraph, none of this is legal advice; the project's ship-gate is a freedom-to-operate review.)

Companion material: the executed notebook garden.ipynb, which measured every claim below, and the eno_render tool's garden_played and garden_idle scenarios, the listening copies. The Max wrapper lands in the TapTools-Max package alongside the rest of the family.

Signal-flow diagram of tap.garden~: notes through a scale quantizer into a 64-event ring, fired at their loop positions into a 16-voice wind-chime pool that sums onto a stereo pair, with a red per-pass path multiplying velocity by decay and brightness by soften back into the ring, and a dashed seeded gardener planting into the ring

Events on a loop instead of audio on a tape — the same recirculation, one level of abstraction up.

Plant and return

note(pitch, velocity) plants: the pitch snaps to the current root and scale at entry, a small wind chime is struck on the next sample, and the event takes a seat at the loop's current position. Every pass, it fires again at velocity × decay, and below floor it retires. The notebook's staircase is the whole contract in one figure: a plant at 0.8 with decay 0.5 returns with its fundamental at exactly half the last, four times over (measured ratios 0.500, 0.500, 0.500, 0.500), then silence, and active_events reads zero. The whole strike fades a shade faster than its fundamental — quieter returns are also duller, because strike hardness couples brightness to velocity.

A rendered waveform showing five returns of one planted note, each half the height of the last, with the measured peak levels labeled and the retirement floor marked

The return staircase: decay 0.5, floor 0.05, and a bloom that knows when it is finished.

That arithmetic is also the stability story. The family's inversion — degradation as the stabilizer — reaches its third form here: a bloom lives exactly ceil(log(floor/velocity) / log(decay)) passes, so the population of live events converges by construction no matter how fast you plant. And beneath the arithmetic sits a hard bound: sixteen chimes in a fixed pool, the quietest stolen when a seventeenth is needed, its envelopes re-aimed rather than reset so a steal glides instead of clicking.

soften — returns get purer, not just quieter

The chime is four decaying mode doublets at the transverse ratios of the chosen material — by default 1 : 2.756 : 5.404 : 8.933, the free-free-tube physics in Fletcher & Rossing — with the upper modes softer, steeper in brightness (b, b², b³), and dying roughly as f² faster, so the fourth mode is the few-millisecond tick of clapper contact and every strike rings down to its fundamental. Each mode pair is split a few cents, the way a real tube's degenerate modes are, so the tail beats slowly instead of decaying like a lab sine. Each pass multiplies the event's brightness by soften, and brightness is the upper modes' level: a bloom collapses toward its fundamental as it recedes, losing its tick first — the tape chapters' generation loss, restated in modes instead of passbands. The notebook measures the mode-two-to-fundamental ratio shrinking by exactly soften every single return, and the pinned tests hold each piece separately: the tick confined to the contact, the tail's beat dipping and returning, soft strikes duller than hard ones, high tubes ringing shorter than low.

The rack: material, flaws, and seats

material swaps what the tubes are made of — a mode, not a fader. At 0 the rack is wind chimes, the free-free tube's 1 : 2.756 : 5.404 : 8.933; at 1 it is a tuned bar, the mallet instrument's double-octave 1 : 4 : 10 : 20 (both tables from Fletcher & Rossing). The table is read at strike time, so every live bloom re-voices at its next return: the notebook measures the second partial's energy moving cleanly from 2.756× to 4× when the material flips.

And the tube is the identity. Each pitch is a physical tube whose imperfections are properties of the tube, not the strike: its upper modes sit a fixed few cents off the ideal ratios (bounded by ±3 cents — the fundamental stays true, because a maker tunes the fundamental and the overtones land where the metal puts them), and it hangs at a fixed seat on the stereo rack, width set by spread (0 collapses to center mono, bitwise identical busses). Both draws come from a stateless hash of the pitch, so the rack is the same rack in every instance and every return of a bloom rings from the same place with the same flaws — the notebook's seat chart is a bar per pitch, and the seed triad below is untouched because no generator is ever consumed for it.

The scale contract

root and scale (chromatic, major, minor, and both pentatonics — plain public-domain scale theory) define where plants may land, and quantization happens at entry: the notebook plants all thirteen chromatic pitches from 60 to 72 into a C major-pentatonic garden and the YIN oracle reads every sounded note on {C, D, E, G, A}. Wrong notes are not discouraged; they are unrepresentable, which is most of why instruments in this family feel effortless to strangers. Because quantization is at entry, changing the scale re-pitches nothing already planted — the field changes for future seeds only.

The gardener

idle_seconds is the patience: that long after your last plant, the wind picks up. The gardener strikes on a calm/gust cycle — gust at 0 is a still day, single strikes spaced about one per pass; raise it and strikes arrive in flurries of up to five neighboring tubes within a fraction of a second, with longer calms between, the average rate holding. The randomness is the family's seeded xorshift64* with the full tr808 contract, pinned as a triad: same seed, bit-identical garden; different seed, a different garden; gardener disabled (idle_seconds 0), the seed cannot matter at all, because the generator is never consumed. This is the library's first randomized event source — step_seq.h proudly promises "no randomness anywhere" — and the seed contract is what lets a generative instrument live in a test suite that demands reproducibility.

Recipes

  • The lobby: defaults, @idle 30. @level 0.4, plant four or five notes, walk away. The garden holds the room indefinitely, bounded.
  • The music box: @decay 0.5 @soften 0.7 @idle 0 @bell 0.005 0.8 1. — no gardener, fast decay: each phrase you play unwinds itself to silence in a few passes, a wind-up toy running down.
  • The endless install: @scale minorpentatonic @root 2 @idle 3. @gust 0.6 @seed 2008 @level 0.35, never touch it again. Same seed next year, same garden — gusts and all.
  • The still day: @gust 0 @idle 10. — no flurries, one unhurried strike at a time, the original music-box gardener.
  • The marimba loft: @material 1 @spread 1. @decay 0.7 @soften 0.8 — tuned bars instead of tubes, the rack thrown wide: drier, woodier blooms that each speak from their own place in the image.
  • Duet: @idle 6. and stay at the keyboard — every silence longer than six seconds, the gardener answers you; every plant of yours resets its patience.

The same machine, in pieces

The four machines inside this one — the entry quantizer, the event ring, the chime rack, the seeded gardener — are objects too: tap.scale, tap.bloom, tap.chime~, tap.gardener. Chained, they are this object bitwise, gardener and all (pinned in tests/garden_test.cpp). The one worth reaching for on its own is tap.bloom: separated from the chime it recirculates notes and has no opinion about what sounds them, so the principle will drive a sampler or MIDI out just as happily. One difference to know before you patch it — out here the ring runs on Max's scheduler rather than the audio clock, so returns land within a millisecond of the grid instead of exactly on it. See The same machine, in pieces.

When it is not the right tool

  • Melodies with wrong notes in them. Quantization is always on; chromatic passing tones survive only in @scale chromatic, and micro-tonal pitches not at all. This is a fence, and it is the product.
  • Rhythm. Events return on the loop grid, exactly, forever — no swing, no humanization. For patterns as rhythm, tap.808.seq~ is the machine.
  • Any other timbre. Two materials, one chime family, on purpose. It is an instrument, not a polysynth; for synthesis as a playground, patch oscillators.
  • A stereo panner. The image is a rack of fixed seats keyed by pitch — there is no per-strike pan and no motion. For placement as a parameter, pan the object's output.

Checkpoint

Notes become events; events recirculate on a loop, quieter by decay and purer by soften each pass, retiring below floor; a sixteen-chime pool bounds the sound and a sixty-four-seat ring bounds the score, oldest bloom yielding first. Two materials share the rack, every tube keeps its own flaws and its own stereo seat, the scale makes wrong notes unrepresentable, and a seeded gardener keeps the piece alive exactly as long as you neglect it. Every number above lives twice: as an executed cell in garden.ipynb and as a pinned scenario in tests/garden_test.cpp, which CI runs on every push.

The same machine, in pieces

The two chapters before this one describe instruments you switch on and walk away from. tap.airport~ turns seven loops; tap.garden~ tends itself. That is the right shape for what they do, and neither is going anywhere.

But both were monoliths by accident rather than by design. Open airport.h and there was never a loop bank doing loop-bank things — there was an array of eight identical lanes and a summing loop. Open garden.h and there was a quantizer, an event ring, a chime rack, and a seeded gardener, wired together by a class that did nothing else. The parts were already there. Nothing outside the monolith could reach one.

So they were promoted. The lanes and the parts are objects now, and the block diagrams at the top of the last two chapters are patchable:

ObjectWhat it isWas
tap.reel~one free-running tape loopa lane of tap.airport~
tap.chime~the sixteen-bell wind-chime rackthe voice pool of tap.garden~
tap.chime.voices~the same rack, one bell per outlet—
tap.bloomthe event ring — plant, return, fade, retirethe recirculation of tap.garden~
tap.scalesnap a pitch to a root and scalethe entry quantizer
tap.gardenerthe idle wind, seededthe self-seeding half
tap.periodwhen a set of loops realignsthe bank's period message

The monoliths remain exactly what they were. This is additive: the same kernel classes, reached two ways.

"The patch is the object" is a measurement, not a slogan

It would be easy to say that three tap.reel~ summed are a tap.airport~ and leave it there. The house rule is that claims of that kind get measured, so this one is pinned in CI like any performance number.

The scenario "standalone lanes summed are the bank, bitwise" in tests/airport_test.cpp configures a three-lane bank and three standalone lanes identically — incommensurate lengths, both exact pan endpoints and one interior pan, one shaded darken corner and one bypassed — drives both through the same staggered punch-in schedule for two seconds, and requires the two stereo outputs to be equal to the bit, not to a tolerance. Nudging one lane's level by 1e-12 fails it.

The garden's version, "the bed is exactly its components wired together, bitwise" in tests/garden_test.cpp, does the same across twenty seconds with the seeded gardener running — which puts the order of random draws under test too, since that is the part a careless split moves without anyone noticing.

Bitwise is available here because the objects' own promises are already bitwise: transparent playback of a frozen loop, exact pan endpoints, a darken stage that is genuinely bypassed at the band ceiling. A decomposition can be held to the same standard the object is.

What you get for patching it

For the airport, four things the monolith cannot give you:

  • An insert on one loop. A filter, a reverse, a tap.discreet~ for tape breath on one phrase and not the others. Inside the bank every loop gets the same treatment, which is to say none.
  • A varispeed on one reel — the bank has one shared clock by construction.
  • More than eight loops. Eight was a number, not a principle.
  • Tape you actually use. The bank buys all eight worst-case reels at DSP start whether you use them or not: about 92 MB of double tape at the 30-second default. Three tap.reel~ buy three, about 11 MB each.

The rack has a second form worth knowing about. tap.chime.voices~ is the same sixteen bells with each one on its own outlet, carrying its tube dry — before the seat in the stereo image. Filter one voice and you are filtering whichever bell happens to be in that slot, not the rack; it is a different instrument, and there is no way to ask tap.chime~ for it. Because the pool reassigns bells as it steals, a slot is not a pitch, so the object will tell you which tube it is holding and what seat it would have been given. Sum the sixteen back through those seats and you have tap.chime~ again, bitwise — pinned by "the per-voice taps summed through their seats are the stereo rack", across twenty strikes, four more than the pool holds, so stealing is under test too.

It is a separate object rather than a switch because outlet count is fixed when a Min object is built, and it is sixteen discrete outlets rather than one multichannel outlet because min-api's mc support is inlet-side only: it sets Z_MC_INLETS and offers no multichanneloutputs, which is what Max requires before an external may declare a variable-channel mc outlet. That is a limitation of the wrapper we have, not of the idea.

For the garden, the interesting one is tap.bloom. Separated from the chime it turns out to be the most portable idea in the family, because it recirculates notes and has no opinion about what sounds them. Point it at makenote, at a sampler, at MIDI out, and Eno's principle — a touch becomes a note, the note returns a little quieter each pass until it is gone — drives an instrument that has nothing to do with wind chimes.

Splitting also made two promises directly testable that were previously only reachable through audio. The ring's arithmetic is now countable with no envelope tail in the way: "the ring's convergence theorem is exact when nothing sounds it" checks four different velocity/decay/floor triples against ceil(log(floor/velocity)/log(decay)) exactly. And the rack's allocator can be watched directly — "the rack fills idle bells first, then steals the quietest" fills the pool, strikes a seventeenth tube, and measures that the faint tube lost its partial while a loud one kept its own.

Where the seams show

Three honest costs, none of them hidden.

The garden's patch is not sample-accurate. tap.bloom and tap.gardener run on Max's scheduler rather than the audio clock, so a return lands within an @interval tick — a millisecond by default — instead of exactly on the sample. Inside tap.garden~ the same ring is sample-accurate. At loop lengths measured in seconds nobody will hear the difference, but it is a difference, and it is why the garden's null test lives in the kernel where both sides can share one clock, and why there is deliberately no in-Max null test for it. Asserting a null that cannot hold would be worse than not asserting one.

Voice stealing had to stay in the kernel. The obvious Max answer to a sixteen-voice rack is one voice in a poly~. That answer is wrong twice: poly~ steals round-robin, which loses the whole point — this rack steals the quietest bell and re-aims it, so its phases keep free-running and its seat glides rather than clicking — and poly~ does not exist off Max, while the kernel is meant to run anywhere. So tap.chime~ is the whole rack, and its polyphony is its own.

A bell reads silent until it has been processed once. The allocator asks each bell for its level, and a bell that has been struck but not yet processed still reports zero. Strikes issued in the same sample therefore land on the same voice instead of spreading across the pool. Inside the bed this only happens when two blooms share a loop position; it is pre-existing behaviour, and it is documented in the rack scenario rather than fixed, because fixing it would change how the object sounds.

The one thing the airport decomposition looked like it would lose is composite_period — the report of when the whole system realigns, which needs every length at once and so has nowhere to live inside a single reel. That arithmetic came out of the bank as a free function instead, and tap.period is it: hand it the lengths and it answers in seconds, inf included. The detail that makes it trustworthy rather than merely convenient is that it shares the reel's seconds-to-samples quantization rather than copying it. The lcm is over sample counts, and lengths that look commensurate written down are usually nothing of the kind once rounded to samples — 0.5 and 0.625 seconds realign at 2.5 s, while the terminal recipe's seven lengths leave the 64-bit range entirely. Both are pinned in tests/airport_test.cpp.

Checkpoint

Seven objects, almost no new DSP: the same kernel classes the monoliths hold, given names and inlets. Three tap.reel~ summed are a tap.airport~ bitwise; tap.gardener into tap.scale into tap.bloom into tap.chime~ is a tap.garden~ bitwise, gardener and all. Both identities are pinned scenarios in tests/airport_test.cpp and tests/garden_test.cpp, which CI runs on every push, and the airport's is checked again against the real externals loaded in Max by runtime-tests/patchers/tap.reel~-is-airport.maxtest.maxpat. The monoliths still do what they did — every scenario that pinned them before the split passes unchanged after it.

Four heads and a motor

The last chapter's tap.discreet~ is a machine you set up and walk away from. tap.tapecho~ is one you keep your hands on. Same spool of tape, same worn return path, same family — but where the Eno objects are systems that run without you, this one is an instrument, and every parameter on it is a hand on the machine. That is the thread through this part of the book: these are the objects you ride.

What it recreates is the tape echo of the Copicat / Space Echo school: one record head, a span of moving tape, several playback heads at fixed positions along it, and a path from the heads back to the record head. Ed O'Brien's Copicat is the reason it is here. It is a recreation of the topology, not a circuit model of any one unit — the tape path itself is the same published tape-echo modeling literature tap.discreet~ already stands on (Arnardóttir, Abel, and Smith's AES model of the Echoplex, and Välimäki et al.'s tape-echo work), and no head spacing, filter curve, or trim value in this object is claimed as measured from a real machine.

Companion material: the executed notebook tapecho.ipynb, which measured every number below, and the radiohead_render tool, whose tapecho_heads, tapecho_three_head, tapecho_selfosc, and tapecho_varispeed scenarios are the listening copies — all four performed, with the controls moving while they render, because static settings tell you almost nothing about this object.

span — the motor

span is the delay of a head sitting at the far end of the tape path, and every other head sits at span times its own ratio. So span is not "the delay time" of one echo; it is the motor speed, and moving it moves the whole layout together.

An impulse into a 400 ms span with four heads, returning once at each of 100, 200, 300 and 400 ms on the dotted head positions

One impulse, four heads. The returns land exactly on span × ratio.

Moving the motor while audio runs is a tape-speed change, which means it bends pitch on the way — the same doppler contract as tap.discreet~, for the same reason: the heads are physically moving relative to the tape. smooth sets how long the motor takes to change speed, and therefore how deep the bend is. There is no crossfading "digital" mode. If a pitch bend on a delay-time change would ruin the patch, reach for tap.delay~.

heads, ratios, levels, pans — the layout

Four heads by default, evenly spaced at 0.25, 0.5, 0.75 and 1.0 of the span. That spacing is nominal — chosen because it is neutral and audibly a tape echo — and every ratio is freely settable underneath, which is how you build a three-head Copicat-style layout:

heads 3, ratios 0.333 0.667 1.

levels is per-head gain and pans places each head in the stereo field (equal-power, with exact endpoints: a hard-panned head is bitwise absent from the far bus). One thing to know: a head's level is also its send into the regeneration path, as the head selector on the real machines is. Turn a head down and you are turning down both what you hear from it and what it feeds back.

regen, drive, darken — past unity, on purpose

Here is where this object parts company with everything else in the house. tap.delay~ caps feedback at 0.99 so the loop is always contractive. tap.discreet~ reaches exactly 1.0 because the wear path is the stabilizer. tap.tapecho~ goes past 1.0 — up to 1.5 — into deliberate sound-on-sound self-oscillation, the howl you reach for this machine to get.

It stays bounded because the saturator does. drive is record-head saturation, and its output can never exceed 1/drive no matter what the loop accumulates, so the tape is bounded by the input plus regen/drive whatever the loop gain. The measurement is the point:

Measured peak output against drive at regeneration 1.4, sitting below the analytic ceiling at every drive

Regeneration at 1.4 — well past unity — plateaus under the saturator's ceiling at every drive.

Because that bound only exists while the saturator is engaged, the effective regeneration is capped back to 1.0 whenever drive is 0 — and the cap is applied per sample, so dropping drive mid-howl lands the loop rather than letting it run away. The attribute keeps its value and takes effect again when drive returns. Twelve seconds of ring at regen 1.4 measures a growth ratio of 1.007 between the two late windows: it plateaus, it does not climb.

darken is the per-pass corner. Every trip through the regeneration path runs through a one-pole lowpass, so the repeats lose treble generation by generation — measured at 0.2915 of a 6 kHz tone per pass against 0.2920 predicted, and 0.8895 of a 300 Hz tone against 0.8898. Riding darken while the loop howls is a performance control, not a set-up step; it is what turns a howl into a swell and back.

wow and flutter — one motor, one path

The transport is the family's deterministic pair of sines, and one motor moves the whole tape path, so a speed error displaces every head together. The pitch math is checkable in closed form: depth times 2π times rate is the peak deviation, so 2 ms at 0.5 Hz predicts ±10.88 cents and the notebook's pitch track measures 10.91. Two renders of the same settings are bit-identical — periodic and deterministic by design, with stochastic capstan drift a documented non-goal, because bit-exact renders are what let the oracle test exist at all. Set both depths to 0 for a still machine.

The one that is not a knob

With the tape path neutralized — no transport error, no regeneration — a one-head echo is bitwise tap.multitap~ with one tap. Same Hermite read, same fractional position, same equal-power pan law. That is not a curiosity; it is the whole design claim, measured: this object is composition over the shared tape machinery rather than a second implementation of it, and tape_loop.h needed no changes at all to serve a topology it was not written for. The appendix has the derivation.

Recipes

  • The Copicat: heads 3, ratios 0.333 0.667 1. with @span 390 @regen 0.6 @drive 0.9 @darken 2600 @wow 0.9 0.9 @mix 50. Heads down the middle, a tired transport, repeats that thicken as they recirculate.
  • A wide slap: four heads, pans -0.7 0.5 -0.35 0.8, @span 480 @regen 0.45 @drive 0.4 @mix 45. The layout does the widening; no chorus needed.
  • Sound-on-sound: @drive 0.7 @regen 1.35, then bring @input to 0 and take your hands off. Ride @darken down to 1400 while it howls, then @regen 0.55 to bring it home. clear is the emergency stop.
  • The dive: @smooth 3000, then @span 200 → @span 900. Three seconds of tape slowing down, with everything already on the tape bending with it.

When it is not the right tool

  • Tempo-locked delays. Span changes bend pitch by design and there is no sync. tap.delay~ is the clean line.
  • A wash you set and leave. That is tap.discreet~, one chapter back — same machinery, opposite posture.
  • Independent free-running loops. One motor moves every head here. For loops that drift against each other, tap.airport~.
  • Clean repeats. Wear is always in the regeneration path; drive 0 removes the saturation, not the darkening.

Checkpoint

A motor and up to four heads along one tape path; the motor moves them together and bends pitch doing it. Regeneration goes past unity into self-oscillation, bounded by the saturator rather than a gain cap, and capped back to 1.0 the moment drive leaves. The transport is two deterministic sines measured in cents. And with the tape path neutral the whole object collapses, bitwise, into a delay this library already had — which is how you know it is composition and not a rewrite. Every number above lives twice: as an executed cell in tapecho.ipynb and as a pinned scenario in tests/tapecho_test.cpp, which CI runs on every push.

The part that comes apart

There is a moment at the end of "Go To Sleep" where the guitar stops being a guitar. It does not fade, it does not filter — it starts eating itself, firing fragments of the bar you just heard in an order nobody played. That sound came out of a Max patch Jonny Greenwood built and performs live. So tap.stammer~ has an odd position in this package: it is a Max stutter object, in a Max package, for a technique that was invented in Max.

None of which means anything was copied. This is an original design in the brassage tradition (Roads, Microsound) — a continuously recorded buffer, a rhythmic grid, and dice. What the band's rig contributes is the knowledge of what the object is for, which turns out to be the hard part of designing one.

Companion material: the executed notebook stammer.ipynb, and the radiohead_render scenarios stammer_grid (the dials held still so the mechanism is audible), stammer_disintegrate (forty seconds of the performance the object exists for), and stammer_two_seeds.

How it works, in one paragraph

The input is captured continuously into the last few seconds of history. On a step grid, if the machine is idle, it rolls: with probability density it grabs the material that just went past, chops it to step divided by something between 1 and divisions, and plays it back between 1 and repeats times, each pass with a reverse chance of running backwards. If jump is open it may reach further back than the bar just played. While a slice fires you hear the slice; when nothing is firing you hear the input, untouched.

density and repeats — they are not the same dial

This is the one thing worth internalizing before you patch it. density is how often the machine grabs; repeats is how long it holds on once it has. A slice in flight is never interrupted, so repeats is what actually decides how busy the machine is — and once trains start overlapping, raising density stops doing anything at all.

Four runs showing when a slice is in flight: raising repeats from 1 to 6 fills the timeline far more than raising density from 0.3 to 0.9

Measured off the object's own playing flag, 100 grid points per run. Repeats is the hold.

At density 0.3, going from 1 repeat to 6 takes the machine from 41% busy to 76%. At density 0.9 it is already 90% busy with a single repeat, and 96% with six — the ceiling, where the dial has run out of room.

divisions, reverse, jump — the character

divisions is how finely the grid may be chopped: at 1 you get whole-step slices, at 8 the machine may cut down to eighths of a step. Because the divisor is drawn per slice, a high setting gives you a mixture of lengths, not uniformly short ones — which is what keeps it sounding played rather than gated.

reverse is drawn per repeat rather than per slice, so a single train can stagger forwards and back. jump is the reach: at 0 the machine only ever replays the material immediately past (the classic stutter), and opening it lets slices come from seconds ago, so the part starts quoting itself out of order. That is the setting that turns "stuttering" into "disintegrating".

fade is the anti-click — a raised-sine flank on each repeat, exactly zero at the edges and exactly unity across the plateau. Repeats are sequential rather than overlapped, so every junction dips to zero. That is deliberate: the dip is the articulation of a stutter, and a crossfade there would smear the thing you want to hear.

seed — the dial that is a contract

Every draw — fire, division, repeat count, reach-back, and the per-repeat coin for reverse — comes from a seeded generator in a fixed order. So the same seed and the same moves give the same render, bit for bit. That is not a nicety; it means a take you liked is recoverable, two instances on different seeds decorrelate instead of moving in lockstep, and the tests can assert bitwise equality. A different seed is a genuinely different performance: 89% of samples change.

And at density 0 the dice are never rolled at all — so the seed provably cannot matter, and the object is a bitwise bypass at any mix. Switched off, this is not "nearly transparent", it is your input. clear erases the capture, drops the slice in flight, and rewinds the seeded stream, so the same seed replays from there.

The material contract

The header says this object wants transient material, and that on a sustained pad a stutter is barely a tremolo. That reads like taste. It is not — it is a property of the material, and it is measurable.

Slices of a sustained sine measure 1.000 alike; slices of a plucked phrase measure 0.286

How alike two arbitrary slices of the material are. Re-ordering interchangeable things does nothing.

Every slice of a steady sine looks like every other slice, so shuffling them changes almost nothing you can hear. Slices of a played phrase are all different, so shuffling them is the entire effect. Feed this object drums, plucked or struck strings, consonants — anything whose interest is in when things happen. It re-articulates rhythm that is already in the sound; it cannot invent rhythm that is not.

Recipes

  • A grid you can hear: @step 250 @density 0.55 @divisions 4 @repeats 4 @reverse 0.2. The mechanism, plainly, over a played part.
  • The disintegration: start at @density 0.2 @divisions 1 @repeats 1 and walk over thirty seconds to @density 0.9 @divisions 8 @repeats 10 @reverse 0.6, tightening @step from 250 to 120 as you go. Then open @jump 1500 and the machine starts quoting the wrong bar.
  • Vocal chop: @step 125 @density 0.4 @divisions 2 @repeats 3 @fade 6. Consonants are transients; the longer flank keeps it from sounding digital.
  • Two of them: the same settings on two instances with different seeds, panned apart. They decorrelate by construction — that is what the seed contract buys you.

When it is not the right tool

  • Sustained material. See above; it is measured. A tremolo or a gate will do more for a pad.
  • Pitched mangling. Slices play at ±1 rate only — there is no pitch shift and no varispeed here. tap.shift~ transposes; tap.pitchaccum~ spirals.
  • Exact, notated rhythms. The grid is regular but the dice are dice. tap.808.seq~ sequences; this improvises.
  • Very long repeat trains. A slice reads from the ring, not a private copy, so a train longer than the captured history will start reading fresher material as the write head laps it. Size the object argument to the longest train you intend to fire.

Checkpoint

Capture everything, then on a grid roll dice and re-fire what just went past. density grabs, repeats holds — and holding is what fills the timeline. divisions, reverse and jump are the character, and jump is the one that turns a stutter into a disintegration. The seed is a real contract: same seed, same performance, bit for bit; at density 0, a bitwise bypass. And the material contract is measured rather than asserted, which is the honest way to tell you what to feed it. Every number above lives twice: as an executed cell in stammer.ipynb and as a pinned scenario in tests/stammer_test.cpp.

The dirt with two stages

Two objects in this library make things dirty and they are not competing. tap.overdrive~ is a feedback soft-clipper chasing the Tube Screamer lineage — the nonlinearity sits inside a loop with a lowpass, so the bass stays clean and the mids break up first. tap.fuzz~ is the other school: two clipping stages one after the other, and a tone section that scoops the middle out. It is the OK Computer-era sound — the dirt on Paranoid Android and My Iron Lung — and it belongs in this part of the book because, like the tape echo and the stutter, the interesting settings are the ones you arrive at by moving something.

The method is not invented here. It is the simplified cascade of Yeh, Abel and Smith's DAFx-07 paper on distortion and overdrive pedals: conditioning filter → memoryless nonlinearity → equalization filter, twice. That paper also supplies the licence for the central shortcut. A real diode limiter is not a static curve at all — it is a lowpass whose pole moves with the voltage across it, and solving that honestly is expensive. Approximating it as a fixed curve between fixed filters is defended there, and measured against real pedals.

What this object is not is a model of a specific pedal. No resistor, capacitor or corner frequency in it is claimed as measured from a unit, and the control names follow the layout that class of pedal conventionally carries rather than asserting what any particular one does.

Companion material: the executed notebook fuzz.ipynb, and the radiohead_render scenarios fuzz_gain_sweep, fuzz_tone and fuzz_edge_and_bite.

One curve, two knees

Both stages share a single clipping family — tanh(kx)/tanh(k) — normalized so that full scale in is full scale out at every knee. That normalization is what lets the knee be a character control instead of a hidden volume control.

Four clipping curves at knees 0.5, 1.6, 2.0 and 12, all passing through the same full-scale point, the sharpest approaching a hard corner

The knee sharpens the corner without moving the ceiling.

The first stage takes a soft knee and most of the gain (the op-amp-ish stage); the second takes a harder one at unity (the shunt limiter). edge sweeps the second stage's knee from a gentle limiter toward something close to a hard corner.

gain — and why the floor is below unity

The knob sweeps the first stage's drive. Its floor sits below unity deliberately, and the reason is the most useful thing in this chapter if you ever build a cascade of your own.

The tanh family's small-signal gain is k/tanh(k) — greater than one, and growing with the knee. Put a fixed ×2.2 in front of a knee-3 curve and the second stage sees an effective ×6.6, which means it is fully clipped before the gain knob leaves zero. That is exactly what the first version of this kernel did. It sounded like a distortion at every setting, which is precisely why listening did not catch it and a measurement did.

Two panels: harmonic content rising steeply with the gain knob, and even-to-odd harmonic ratio rising with asymmetry

Left: the gain knob after retuning — harmonic content sweeps 0.010 to 0.358. Right: asymmetry is what makes even harmonics.

asymmetry — the even harmonics

A symmetric curve is an odd function, so it can only make odd harmonics. DAFx-07 points out that a real op-amp stage clips lopsided, and that this is where a pedal's even-order content comes from — which is the whole reason this control exists. Turn it up and the even/odd ratio climbs from essentially zero to about 0.55.

It costs no DC. The bias is applied inside the curve and corrected at the stage output, so however lopsided the setting, silence in is exactly silence out — no pedestal, no thump when you stop playing.

bass, treble, contrast — the voicing

Three linear filters entirely outside the nonlinearity: a low shelf, a high shelf, and a mid scoop whose depth is contrast. On this class of pedal the voicing section is most of the identity — the scoop is the sound people mean when they describe it — so it is a first-class part of the object rather than an afterthought bolted on at the end.

oversample — and a default that was wrong twice

A static curve makes harmonics without limit, so anything above Nyquist folds back. The clipper pair therefore runs oversampled. Everything about this control has been re-measured, because the first two conclusions drawn from it were wrong, and wrong in the same way.

First, the anti-alias filter here is 8th order, where the rest of the house uses 4th. Measured in this kernel the 4th-order pair is not steep enough — alias energy at 4× came out worse than at 2×.

Second, the chain is a cascade of 2× stages — one doubling, one filter, repeated — rather than a single zero-stuff by the whole factor. That is what finally made more oversampling mean less aliasing. The single-stage chain left N−1 images for one filter to suppress at a corner that got tighter with every doubling, and the residue intermodulated in the clipper into exactly the non-harmonic junk the probe measures. Cascading removes the reversal outright, and where 4× and 8× used to be merely adequate they are now two to four orders of magnitude cleaner. It costs about 5 % more CPU at 8×.

Third — and this is the part worth taking away — the old default came from a single test tone. Every number in the original write-up was measured at 3733 Hz, and 2× happens to look best there. Swept across tones, 2× collapses above about 6 kHz; at 10.5 kHz it is worse than not oversampling at all, because the clipper's low harmonics already exceed the base Nyquist and one doubling does not move them out of the way.

input tone1×2×4×8×
3733 Hz1.2e-13.0e-52.1e-52.2e-5
5171 Hz1.5e-13.3e-42.0e-71.9e-7
6421 Hz8.5e-23.2e-23.9e-73.6e-7
8123 Hz9.0e-27.8e-21.1e-32.0e-5
10499 Hz1.5e-11.7e-11.2e-51.6e-6

So: 4× is the default. More is never worse now, and 4× is indistinguishable from 8× below about 7.5 kHz. Above that, harmonics start folding inside the 4× band before decimation — 8123 Hz in the table is that happening — and 8× is worth the extra 1.4 % of a core.

Use 2× only if you have measured your own material and it holds up there. It is kept because it is cheap and because on a bass-heavy source it is fine, not because it is good.

Recipes

  • Edge of breakup: @gain 0.3 @edge 0.2 @contrast 0. @bass 0.. Barely dirty; a boost with attitude.
  • The scoop: @gain 0.8 @edge 0.6 @contrast 1. @bass 0.4 @treble 0.2. The sound the control is named for.
  • Lopsided and mean: @gain 0.9 @edge 1. @asymmetry 0.7 @oversample 8. Hard knee plus even harmonics, and 8× because a hard knee on a bright source is exactly where the top octave folds.
  • Into the echo: tap.fuzz~ → tap.tapecho~ with the echo's @drive low. Two saturators in series get muddy fast; let the pedal be the dirt and the tape be the space.

When it is not the right tool

  • Amp-like breakup. tap.overdrive~ keeps the bass clean by design; this object does not, and hard settings will get woolly on a bass-heavy source.
  • Subtle warmth. Two stages is a lot of stages. At low gain this is a clean boost with a tone stack, which is fine, but tap.overdrive~ is the better instrument for gentle.
  • A specific pedal. This is that pedal's class. If you need a named unit, this is not it and does not pretend to be.

Checkpoint

One clipping family with a knee control, cascaded twice, into a voicing section that scoops the middle. The gain knob's floor is below unity because small-signal gain compounds through a cascade — a lesson that cost this kernel one wrong first draft. asymmetry is the even-harmonic control and costs no DC. And the oversample setting is a measurement twice corrected: cascaded 2× stages, because a single zero-stuff by N was what made bigger measure worse — and a default of 4× rather than 2×, because the old default had been generalized from one test tone. Every number here lives twice, as a cell in fuzz.ipynb and as a pinned scenario in tests/fuzz_test.cpp.

Two hands on the same tape

tap.stammer~ and tap.scrub~ record the same way. Both keep a rolling tape of what just went past — the same capture, literally the same code, not a second copy of it — and both put a read pattern on top of it. The stutter's pattern is a slicer with dice. The scrub's is a pad you drag.

That is the whole difference, and it is the difference between a machine that decides and a machine you play. The stammer is a die you load; the scrub is a surface you push around, in the Kaoss-pad school of instruments where an XY surface over live capture is the entire interface. It belongs in this part of the book for the same reason the tape echo does: nobody sets this object up and walks away from it.

It is an original design in the granular / brassage tradition (Roads, Microsound, MIT Press 2001) — not a port, and not a reconstruction of any product. No preset, timing or parameter value in it came from a piece of hardware.

Companion material: the executed notebook scrub.ipynb, the pinned scenarios in tests/scrub_test.cpp, and the radiohead_render scenes scrub_gesture and scrub_freeze.

The two axes are actually two axes

position is how far back the playhead sits, as a lag in milliseconds behind the live edge. pitch is transposition in semitones. On tape those would be the same knob — moving the head is the pitch change — and the whole point of doing this with grains is that here they are not. Hold the position and sweep the pitch and the material transposes without going anywhere. Sweep the position at a fixed pitch and you rake through the last few seconds without the tape rising or falling.

drift is the third one, and it is the playhead's own motion through the tape in playback-rate units: 1. runs forward at the speed the recorder is writing, 0. holds station, negative runs backwards. Set @drift 1. and let go of the position and the scrub is a delay; set @drift 0. and it is a freeze that you can still transpose.

The identity underneath it

Grains are Hann-windowed and fired every size / overlap samples. Hann overlap-adds to exactly 1 at that hop, so with the pitch at unity, spray at zero and the position held on a whole sample, the scrub is the input, delayed, to floating point — 4.4e-16 in the pinned test.

Two panels: the scrub's output lying exactly on top of the input delayed by 480 samples, and the window sum for overlaps 1, 2 and 4 — flat at 1 for 2 and 4, dipping to zero for 1

Left: held still at unity, the object is a delay and nothing else. Right: the window sum that makes it one.

This matters more than it sounds. Everything else the object does is a departure from a plain delay, and a departure is only trustworthy if you know the thing it departs from is exact. When the position drags, when the pitch moves, when spray scatters the origins — those are the object working. If the still case were approximate, you could not tell them apart from noise.

overlap 1 leaves gaps between grains, which is the dipping curve in that figure. That is a chopped, gated texture rather than a defect, and it is worth having; it is just not the setting the null lives at.

What transposing costs, honestly

Reading tape at a rate the write head does not share means the read pointer drifts away from where the position says it is, and it has to be pulled back or the position stops meaning anything. Every pull-back is a splice between two grains reading material a little apart.

What that costs is not the pitch. Swept over seven fundamentals and seven intervals, 98.8 % of a perfect shifter's energy lands within ±15 Hz of the transposed pitch — worst case 91.7 %. The note is where you asked for it.

What it costs is concentration. The band holds a narrow comb rather than one clean line: 92.0 % as concentrated as a clean shift, 75.0 % at its worst.

Two curves against pitch from −12 to +19 semitones: energy at the transposed pitch staying near 1, and concentration dipping to about 0.75

The pitch goes where you put it. What the splices take is focus.

Audibly that is a warble, and it is the classic single-delay-line pitch-shifting artifact rather than anything peculiar to this kernel. If you want the warble gone, spray trades the comb for a broadband smear, which some material prefers. If you want a clean shift, this is the wrong object — see below.

freeze, and what it does not stop

freeze stops the recorder. The playhead keeps going, so the position now addresses fixed tape and the grains loop the same window: a granular hold you can still scrub, transpose and drift through. It does not stop time inside a grain — a grain in flight when freeze engages was already scheduled, and it finishes.

spray and seed

spray scatters each grain's origin randomly back from the position. At exactly 0 the dice are never rolled, so the seed provably cannot matter — the same contract tap.garden~ and tap.stammer~ carry, pinned by the same kind of test. With spray up, the same seed and the same moves give the same render bit for bit, and two instances decorrelate by seed alone.

Recipes

  • The pad: @drift 0. @size 80 @overlap 2 @mix 100, then ride position with a signal. The default instrument.
  • Granular freeze: @freeze 1 @drift 0. @size 120 @spray 40. Hold, then move pitch for a chord that was never played.
  • Backwards tape: @drift -1. @pitch 0. — the playhead walking against the recorder.
  • Chopped: @overlap 1 @size 40. Gaps between grains, on purpose.
  • Into the diffuseur: tap.scrub~ → tap.palme~ with the palme's @mix around 40. The strings sustain what the scrub chops.

When it is not the right tool

  • Clean transposition. The warble above is inherent to the method. tap.shift~ and tap.pitchaccum~ are the objects built for that job — with one caveat worth stating plainly: measured on this same sweep, tap.pitchaccum~ retained mean 0.907 of band energy against the scrub's 0.988, worst 0.633 against 0.917. Its two-tap crossfade also puts its strongest spectral line a few hertz beside the intended pitch, which is filed as issue #33. None of that makes it the wrong object — it is a shimmer, and shimmer is what those sidebands are — but audition before you assume it is the transparent one.
  • A tidy delay. tap.delay~ and tap.tapecho~ cost far less and do not window anything.
  • Slicing to a grid. That is tap.stammer~, on the same tape.

Checkpoint

One tape shared with the stutter, one grain scheduler on top of it, and two axes that stay independent because grains let them. The still case is a bit-exact delay, which is what makes every departure from it legible. Transposing warbles, and the warble is measured rather than apologized for: the note holds to 98.8 % of a clean shifter's energy, and 92.0 % of its focus. Every number here lives twice, as a cell in scrub.ipynb and as a pinned scenario in tests/scrub_test.cpp.

Loudspeakers you can play

The Ondes Martenot does not have a loudspeaker. It has a rack of them, and the player chooses. Beyond the plain cabinet — the principal — Maurice Martenot built resonating diffuseurs whose entire job is to colour the signal with a physical body: the métallique (1944–45, patented 1947), a gong driven by a motor transducer, and the palme (1949–50), an electromagnet driving twelve metal strings stretched on a soundboard.

tap.metallique~ and tap.palme~ are those two, and they ship as standalone effects rather than as something hidden inside tap.ondes~, because the interesting thing about a resonating loudspeaker is that it does not care what you put through it. A guitar into the palme is not what Martenot had in mind and it is the best reason to have the object.

Najnudel, Hélie, Roze and Boutin (IEEE/ACM TASLP 28, 2020) name the diffuseur as the stage that "converts the electrical waveform into sound and in turn modifies its spectral content". Wijnand, Boutin, Jossic and Maniguet (Forum Acusticum 2023) describe the instruments and measure the transducer. Everything below traces to one of those two, or is labelled as a recreation.

Companion material: the executed notebook diffuseur.ipynb, tests/diffuseur_test.cpp, and the radiohead_render scenes metallique_stages and palme_halo.

Driven, not struck

tap.chime~ and tap.garden~ already carry this library's modal machinery — mode ratios, doublet splitting, per-mode decay — and it carries over here intact. What does not carry over is the strike. There is no trigger in either of these objects and no decay envelope. A diffuseur is excited continuously by whatever is going through it and rings at its own rates, which is tap.5comb~'s sustained-resonance situation rather than the chime's.

Practically, that is the difference between an object you fire and an object you feed.

The order is the argument

The electrical signal reaches the transducer first, and the transducer's motion is what excites the body. So the nonlinearity sits upstream of the resonator. Drive the transducer hard and you are pushing a distorted waveform into a gong — which is a different sound from distorting a gong.

That claim is pinned rather than asserted: a null test in the kernel checks that a whole cabinet is bitwise identical to transducer → body wired by hand, and that the reverse wiring differs by 28 % of peak. It is not a subtlety you have to take on faith, and it is not a subtlety you can hear your way past.

The métallique

Eight modes at the free circular plate's transverse ratios — Rayleigh's classical Chladni set at Poisson 0.3, 1 : 1.730 : 2.328 : 3.910 : 4.110 : 6.300 : 6.710 : 7.340 — each split into a slowly beating doublet.

Two panels: eight modes as stems with their ratios labelled, weights summing to 1; and the plate's measured response to a swept drive tone, peaking at each mode

Left: where the modes are. Right: the body answering a sweep, which is how you actually meet it.

pitch places the lowest mode and the rest follow. decay is the fundamental's T60 — long is a drone, short is a plate reverb. tilt decides how much faster the upper modes die than the fundamental, and brightness weights them. The weights sum to exactly 1 and each mode has unit peak gain, which is why there is no limiter on the output and no DC blocker either: the body is bounded by its input, by construction.

The palme

Twelve strings, each a damped delay loop, on one board.

Twelve, not twenty-four. Widely copied build pages say two banks of twelve; the peer-reviewed source says twelve, and this object follows the peer-reviewed source.

Their tuning is not published anywhere found, so it is a control: @tuning 0 lays them out chromatically across an octave from root — a string for every pitch class, so the board answers whatever you play — and @tuning 1 puts the harmonic series on the root, which is a drone that answers one key.

The palme's ringing after a faded drive tone is removed, swept from 100 to 440 Hz, with twelve peaks lining up on the twelve string frequencies

Feed it a tone, take the tone away, measure what is left. Every one of the twelve strings rings at least 4.4× harder at its own pitch than between them.

damping is how fast a string loses its upper partials — low values are felt cloth on the strings. detune scatters the strings against each other by a fixed, deterministic amount in cents, because no two strings on a real board are in perfect relation.

The transducer

Wijnand et al.'s point about the early diffuseurs is that they use a moving-iron driver whose operating principle is inherently nonlinear — Thiele–Small does not describe it — so a diffuseur modelled as a pure resonator is missing a documented stage.

What is modelled here is that principle, not a fit to a measurement. In a moving-iron motor the force follows the square of the gap flux, so with a bias current I₀ and signal i the force carries a term in (I₀ + i)² whose residual i² makes second-harmonic distortion that grows with drive. That is asymmetry: the transducer's own even-harmonic signature, and the only part of these objects that is nonlinear by citation.

saturation is the honest exception. A squared law is expansive and something has to bound it, so there is a soft clipper after it — a modelling necessity, not a measured stage, and its coefficient is a knob rather than a number from a paper. At 0 it is exactly linear.

What these are, and are not

The instruments, their dates, their excitation and their transducer type are peer-reviewed. The modal data is not. No ondes-specific measurement of either body exists in any source obtained, so the plate comes from Fletcher & Rossing's free circular plate and the strings from the harmonic series. Both bodies are therefore recreations of the general physics, not models of Martenot's instruments. Nothing here was fitted to a recording, a measurement, or a photograph.

There is also no radiation model — no directivity, no cabinet, no soundboard resonance of its own. The output is the body's modal response, not a room. And the strings are ideal: a real steel string is stiff and its partials stretch sharp, and that dispersion is not modelled. detune scatters strings against each other, which is a different thing and does not stand in for it.

Recipes

  • A guitar into the palme: tap.palme~ @root 110 @tuning 0 @decay 8 @mix 45. The halo underneath everything you play. The reason these ship standalone.
  • The instrument, assembled: tap.ondes~ → tap.palme~ @mix 60. What Martenot actually had.
  • Gong reverb: tap.metallique~ @pitch 180 @decay 1.5 @tilt 1.2 @mix 35. Short decay turns the body into a plate.
  • A drone you drive: tap.metallique~ @decay 20 @drive 3 @asymmetry 0.5 @saturation 0.4 @mix 100. Hard into the transducer, which is upstream, so it is a distorted waveform ringing a gong rather than a distorted gong.
  • The one to be careful with: tap.palme~ @level — twelve resonant loops add up, and a driven board can be much louder than what went into it.

When it is not the right tool

  • A reverb. These are twelve strings and eight modes. They are pitched, and they will impose their pitches on anything you send.
  • A model of Martenot's own diffuseurs. See above: this is the physics of the general case, and the difference is stated rather than glossed.
  • Clean sustain. tap.5comb~ is the sustained-resonance object without a nonlinear driver in front of it.

Checkpoint

Two loudspeakers with bodies, shipped as effects because a resonating cabinet does not care what drives it. Driven rather than struck, so no trigger and no envelope. The transducer is upstream of the body and a bitwise null test pins that it is — 28 % of peak says the order is audible. Every mode has unit peak gain and the weights sum to 1, so the body needs no limiter. And the bodies are recreations of published physics rather than measurements of Martenot's instruments, which is a limitation stated here and in the header rather than left to be discovered.

The instrument that is not a synthesizer

Three objects in this chapter — tap.ondes~, tap.triode~ and tap.touche~ — and one instrument. The Ondes Martenot, 1928, the thing Messiaen wrote for and Jonny Greenwood plays: a keyboard you can also play with a ribbon on a ring, and a pressure key in the left hand that is the whole dynamic range of the instrument.

The plan for this family assumed it would be an oscillator with waveform switches. It is nothing of the kind, and finding that out changed every decision below.

The Ondes Martenot is heterodyne. Two oscillators run near 80 kHz, one fixed and one moved by the ribbon; they are summed, and the note you hear is the envelope of their beating. Najnudel, Hélie, Roze and Boutin, who modelled instrument No. 169 stage by stage (IEEE/ACM TASLP 28, 2651–2660, 2020), measure those oscillators at about 0.03 % second harmonic even coupled to the rest of the circuit. They are essentially pure sinewaves. Every bit of the instrument's character therefore comes from what happens after them: the demodulator, two valve stages, the intensity key, and the diffuseur.

Companion material: the executed notebooks ondes.ipynb and touche.ipynb, tests/ondes_test.cpp and tests/touche_test.cpp, and the radiohead_render scenes ondes_stages, ondes_ribbon, ondes_diffuseurs and triode_tubes.

The biggest source of harmonics is not a valve

Two oscillators of equal amplitude sum to an envelope of 2|cos|. That is not a sinusoid. Its Fourier series puts the second harmonic 14.0 dB below the fundamental, the third 21.3 dB down and the fourth 26.4 dB down — a substantial harmonic series generated before anything nonlinear touches the signal.

Two panels: the envelope of two summed oscillators at depths 1 and 0.5, and the harmonic levels of |cos| — −14, −21, −26, −30, −33 dB

The demodulator is the instrument's largest single source of harmonics, and it is upstream of every valve.

This is why tap.ondes~ does not synthesize a difference tone. Generating the note as a sinewave and distorting it afterwards would throw away the part of the timbre that arrives for free — and it is an easy mistake to make, because the circuit paper does say the oscillators can be replaced by a sinewave generator. That licence applies to the oscillators, not to the demodulator.

The carrier is not simulated either, and that is not a compromise. For amplitudes 1 and depth the envelope is exactly sqrt(1 + depth² + 2·depth·cos φ), so the 80 kHz disappears from the arithmetic rather than being approximated away. Running the published RC detector on that closed form matches a full heterodyne-plus-diode-plus-RC simulation to within 0.10 dB on every harmonic at every pitch tried.

depth is that second amplitude, and it turns out to be the cheapest real timbre control in the object. At 1 the envelope closes completely and the series is full; below 1 it never closes and the tone thins toward a sinusoid. It is a mismatch between two real oscillators, not an invented knob.

detect — and why the instrument thins as it climbs

The detector is the published one: a triode grid near zero bias conducts on positive half-cycles and charges instantly, and R4 × C21 = 1 MΩ × 200 pF discharges it — a 200 µs time constant, which is @detect 0.2.

That single number carries the instrument's pitch character, because an RC that slow cannot follow a fast envelope back down. Measured here, the second harmonic runs from −14.0 dB at A2 to −19.3 dB at A6, and the level falls 2.0 dB across those five octaves. The ondes gets purer and quieter as it goes up, and it does so for a reason you can point at in a schematic.

The ribbon is linear in semitones

The circuit paper's Eq. 7 gives the variable oscillator's capacitance against ribbon displacement, and what falls out is f = 55 Hz · 2^(d / 12·d₀).

So @ribbon is semitones above A1, not Hz. A hand moving at constant speed makes a constant-rate glissando; nothing quantizes, and nothing should. This is why an ondes glide sounds the way it does, and it is the one place where taking the units from the paper rather than from convention changes how the object feels to play.

tap.touche~ — 50 dB in four and a half millimetres

The intensity key is a graphite-and-mica powder bag working as a rheostat: compress it and the number of conducting bead paths rises, so resistance falls. Messiaen called it the instrument's greatest invention. What the player feels is a well-chosen nonlinear spring.

The curve in this object is not modelled and not fitted. Quartier, Meurisse, Colmars, Frelat and Vaiedelich (Acta Acustica 101(2), 421–428, 2015) measured finger force, key displacement and sound simultaneously on instrument No. 320, and published the boundaries of the six musical nuances across the key's travel. Those seven points are the object, interpolated with monotone cubic segments that pass through every one of them.

The published key curve: 50 dB rising steeply between 4.3 and 8.8 mm, seven measured points on it, a straight line for comparison, and the bottom 4.3 mm shaded silent

Seven measured points, and the shape between them. The straight line is what a fit would have thrown away.

Three things follow, and each is a decision the paper made rather than this object:

  • Position, not force and not velocity. The paper states explicitly that the intensity depends on displacement, and not on the speed of the gesture. A static memoryless map is the finding, not a simplification.
  • 50 dB over about 4.5 mm, from 4.3 mm (the instrument's noise floor) to 8.8 mm. The paper notes most traditional instruments rarely exceed 25 dB of per-note dynamic range.
  • The shape is not a line. Equal 8.3 dB steps take displacement steps of 1.0, 0.6, 0.5, 0.4, 0.5 and 1.5 mm. It steepens through the middle and flattens hard at the top.

And the thing that surprises everyone who patches it: on a 0–1 control, roughly the bottom 45 % of the travel is silent. That is not a dead zone in the object. It is the key's own first phase — the elastic strip bending before it reaches the powder bag — and it is exactly why the instrument can be attacked so sharply, because the useful 50 dB lives in the 4.5 mm right after it.

tap.triode~ — the stage is a citation

The valves are where the rest of the character is, and there was nothing to invent. The circuit paper does not merely mention a tube model: it names the enhanced Norman Koren model (Koren, Glass Audio 8(5), 1996, with Cohen & Hélie's grid-current extension, AES 129, 2010), writes out its equations, and publishes parameter sets fitted to the actual valves in ondes No. 169 in its Table II — 6F5 in the oscillators, 6C5 in the demodulator and preamplifier, 2A3 in the power amplifier — along with each stage's supply voltage, cathode resistor and plate load.

A stage is then the static solution of the load line, which is a memoryless nonlinearity in exactly the DAFx-07 sense tap.fuzz~ uses. Where the fuzz reaches for a tanh, this one solves a valve.

Two panels: the 6C5's plate characteristics with its load line and quiescent point marked, and the transfer curves of the three published stages, all sloping downward

Left: the published operating point, solved. Right: the stages invert, and they are visibly lopsided.

Two properties matter before you patch tap.triode~ on its own:

  • It inverts, as a real common-cathode stage does. That is not cosmetic. The valve's asymmetry acts on whichever side of the waveform reaches its grid, so the sign decides which half gets bent.
  • It is strongly asymmetric. At the demodulator's operating point, equal grid swings either way give plate swings in a 2.17 : 1 ratio. That ratio is where a triode's even harmonics come from.

drive is normalized out of the level — the gain-staging lesson tap.fuzz~ learned the hard way, applied here from the start — so turning it up gets dirtier rather than louder.

drive on the voice, and where it starts from

Harmonic content rising from 0.221 to 0.344 as drive sweeps 0 to 8, with the demodulator's own floor marked, while the fundamental level falls gently

The valves add to a signal that was already rich. The floor is the demodulator's.

The important thing in that figure is the dotted line. At @drive 0. the tone still measures 0.221 of harmonic content, because the demodulator made it. The knob sweeps 0.221 → 0.344, monotonically, without the level running away.

The two controls that are choices

Most of this object is a citation. Two controls are not, and both are labelled as such because both measure as audible.

  • keyplacement — the paper's five stages do not include the intensity key, so where it sits is undetermined. After the valves (the default) it is a clean output law: pressure is level. Before them, pressure drives the valves: soft is clean and hard is dirty. The two differ by about 0.09 of total harmonic content at a half-press.
  • polarity — the two valve stages are coupled through a transformer whose winding sense is not in the source, and the sign decides which side of the waveform the preamplifier's asymmetry acts on. Worth about 0.12.

power, and taking the authors at their word

The 2A3 power stage is off by default, following the paper: they measure almost 5 % second harmonic there, but report its contribution as much less important than the two stages before it, and drop it for real-time.

Measured here, switching it on moves total harmonic content from 0.248 to 0.251 and the second harmonic by 0.1 dB. They were right, which is why it is a switch rather than a deletion.

oversample

The nonlinear chain runs oversampled. Worst non-harmonic energy relative to the fundamental, at 1× / 2× / 4× / 8×:

tone1×2×4×8×
587 Hz−79.3−91.2−104.5−103.8
1175 Hz−65.8−77.2−90.6−92.5
1760 Hz−57.6−70.9−81.1−82.2
2637 Hz−51.1−61.4−71.8−83.8
3520 Hz−45.4−56.8−67.0−74.2

Every doubling is worth about 12 dB up to 4×; past that it is worth 7–12 dB at the top of the range and nothing at the bottom, where the measurement has already bottomed out. Never worse. 4× is the default because that is where the cost stops buying uniformly; 8× is there for anyone playing the top octave hard.

Readers of the tap.fuzz~ chapter will notice this used to be the opposite of what that object measured. That was not a contradiction — it was the clue that fixed the fuzz. The appendix explains how.

What is missing, deliberately

The real instrument has waveform registers — switchable timbres. Their filter shapes are in none of the sources obtained, and inventing them is the one thing this object will not do.

There is also no diffuseur in tap.ondes~, because that is tap.metallique~ and tap.palme~, and patching one after the other is how the instrument works anyway.

Recipes

  • The instrument: tap.ondes~ → tap.palme~ @mix 60. Ribbon and key on signals; that is the whole performance surface.
  • Ribbon on a slider: @ribbon takes a signal, and a line~ from 0 to 36 over four seconds is a three-octave glissando that sounds like one because the law is linear in semitones.
  • The key alone: tap.touche~ on any source. It is a published expressive gain law, and nothing about it is ondes-specific once it is detached.
  • Thin and pure: @depth 0.4 @detect 0.6 @drive 0.. The envelope never closes and the detector smooths what is left.
  • Dirty on hard presses: @keyplacement 1 @drive 4 @polarity -1. Pressure drives the valves.
  • A valve on a guitar: tap.triode~ @tube 2 @stage 2 @drive 6 — the 2A3 power stage, used for something it was never in this instrument for.

When it is not the right tool

  • A subtractive synth. There is no filter, no envelope generator and no waveform selection here. It is one voice with a ribbon and a key.
  • A polyphonic anything. The instrument is monophonic; so is this.
  • A specific recording. The valve parameters are a fit to one instrument's tubes, and tube-to-tube spread in 1930s valves is wide.

Checkpoint

A heterodyne instrument whose oscillators are nearly pure, so the character lives downstream: a demodulator whose 2|cos| envelope makes more harmonics than either valve does, two valve stages that are a published model with published parameters, and a pressure key that is a published measurement interpolated rather than fitted. The ribbon is linear in semitones because Eq. 7 says so. Two controls are choices rather than reconstructions and are labelled as choices. The waveform registers are missing on purpose. Every number here lives twice, as a cell in ondes.ipynb or touche.ipynb and as a pinned scenario in tests/ondes_test.cpp or tests/touche_test.cpp.

Making the machine talk

The vocoder is audio's oldest identity theft: take the shape of one sound and wear it over the body of another. Speech works because your mouth sculpts a moving spectral envelope; a vocoder measures that envelope on one signal (the modulator — usually a voice) and stamps it onto another (the carrier — usually a synth), and the synth talks. tap.vocoder~ is the classic architecture: a 24-band channel vocoder, time-domain, no FFT. This chapter is how to wire it and — mostly — how to choose the two signals, which is nine tenths of vocoding.

Companion material: the reference page and help patcher in the TapTools-Max package; the kernel's Catch suite pins the structural behavior quoted below.

The machine, in one pass

Two identical banks of 24 bandpass filters, log-spaced from 50 Hz to 12 kHz (RBJ constant-peak biquads — unconditionally stable across the range). The modulator goes through one bank; a per-band envelope follower measures each band's level. The carrier goes through the other bank; each carrier band is multiplied by the matching modulator envelope; the bands are summed. That's the whole machine — which is why its behavior is so predictable:

  • A silent carrier is silence, no matter what the modulator does (pinned by test): the modulator only ever gates; every sample you hear is carrier.
  • Gain is exactly linear (pinned): the vocoder adds no nonlinearity of its own.
  • A silent modulator decays to silence at the follower rate — the vocoder "lets go" of the carrier the way the voice lets go of a word.

The wiring

Modulator in the left inlet, carrier in the right. Getting these backwards is the classic first-patch bug, and it sounds like it: a synth "speaking" your voice is right; your voice weakly filtered by a synth is backwards.

Signal-flow diagram of the vocoder: the modulator through a 24-band filter bank into envelope followers, the carrier through an identical bank, per-band multipliers, and a summed gain stage

Two identical banks meeting at 24 multipliers. Envelopes gate the carrier; modulator audio never reaches the output.

The knobs, one by one

q — intelligibility vs. smoothness

The bandwidth of all 48 filters. Narrow (high q) separates the bands cleanly — crisper consonant detail, more "robot" — but thins the carrier between band centers. Wide (low q) overlaps the bands into a smoother, duller blend. The classic hardware vocoders sat toward smooth; intelligibility came from performance, not q.

response_interval — how fast the mouth moves

The envelope followers' period in ms. Short tracks every consonant — crisp, maximally intelligible, and a little nervous. Long smears syllables into pads — the "choir" setting. This knob is the vocoder's attack and release; 20–50 ms speaks, 200+ ms sings.

gain

Makeup level, since a band-multiplied signal usually lands quieter than either input. Linear, boring, necessary.

sibilance — the built-in s and t budget

The classic channel-vocoder unvoiced path (Dudley's lineage): a seeded internal noise source blended into the carrier of the bands above ~4 kHz, still gated by the modulator's envelopes — so consonants articulate even over a dull carrier, and only when the modulator actually has high-band energy (pinned: a silent carrier with an HF-rich modulator speaks at sibilance 1; a low-only modulator stays quiet). At the default 0 the original silent-carrier contract holds exactly, bit-identical — turning it up deliberately relaxes that contract for the top bands. The noise is deterministic per seed, family doctrine.

mix — the synth under its own robot voice

Equal-power blend of the dry carrier against the vocoded output — the classic parallel move (the pad fades in under itself talking). Endpoints are exact: 100 is bit-identical wet, 0 returns the carrier untouched.

Choosing the two signals (the actual craft)

  • The carrier must have energy where the modulator has bands. The eternal vocoder failure is a dull carrier: a mellow sine pad gives the high bands nothing to gate, and consonants vanish. The house answer is upstairs in this book — a tap.vco~ saw stack (harmonics forever, and the analog section keeps it moving) is a nearly ideal carrier; noise (tap.noise~) blended in restores the s and t sounds that even a saw can't carry.
  • The modulator wants articulation, not fidelity. Overdriven, compressed, even cheap-microphone speech vocodes better — what matters is envelope contrast between bands, not beauty.
  • Nobody said voice. Drums modulating a pad turns the pad into rhythm; a cello modulating noise is a ghost. The machine imposes any moving envelope on any body.

Recipes

  • The talking synth: speech → left; tap.vco~ saw stack + 10 % noise → right; response_interval 30, q middling, and enunciate like you're annoyed.
  • The choir: sustained "aah"s → left; detuned saws → right; response_interval 250. Consonants don't matter; vowels are the chord.
  • Rhythm transfer: a drum loop → left; anything sustained → right; short response_interval. The drums play the pad.

When it is not the right tool

  • Pitch correction or transposition — a channel vocoder never changes the carrier's pitch; it only shades its bands. Pitch is tap.shift~ / tap.pitchaccum~ territory.
  • High-fidelity cross-synthesis. Twenty-four bands is a voice, not a spectrograph; for surgical spectral morphing you want FFT-domain tools (tap.spectra~ is the start of that corridor).
  • Formant preservation while shifting — related, but a different machine: tap.harmony~, which multiplies the voice itself instead of wearing it over a carrier.

Checkpoint

Two matched 24-band banks, 50 Hz–12 kHz: the modulator's per-band envelopes gate the carrier's bands, and everything you hear is carrier. q trades crispness against smoothness, response_interval is the mouth's speed, and the craft is almost entirely in feeding it a bright, busy carrier and an articulate modulator. The machine is simple; the casting is everything.

A gate for every bin

A noise gate is a bouncer with one rule: too quiet, you don't get in. Useful, but blunt — when the signal plays, all the noise under it walks in too, and when the signal stops, the gate slams on room tone. tap.nr~ hires a thousand bouncers instead: it transforms each STFT frame and applies the threshold per frequency bin, so the quiet bins between your signal's partials close while the loud ones stay open. Hiss disappears from the gaps in the spectrum, not just the gaps in time. This chapter is the two knobs, the two costs, and the one artifact to listen for.

Companion material: the reference page and help patcher in the TapTools-Max package; the kernel's Catch suite pins the reconstruction claims below.

The contract: transparent until it isn't

The object runs its own STFT — Hann window, 4× overlap, COLA-normalized overlap-add — and the engineering contract is pinned by test: with the gate open, the output reconstructs the input exactly (below 10⁻⁶), delayed by one FFT frame. Whatever tap.nr~ does to your sound, it is doing it on purpose with threshold and slope; the machinery itself is transparent. Also pinned: a tone below threshold is strongly attenuated; a tone above passes untouched.

Diagram of the STFT scaffold both spectral objects share: input ring, analysis window, FFT, the pluggable spectral operation, IFFT, synthesis window, and COLA-normalized overlap-add

The pump this object runs on — tap.nr~ is this scaffold with a per-bin downward expander in the middle.

The knobs, one by one

threshold — where quiet begins

The per-bin level (linear amplitude) below which a bin is attenuated. The craft: set it between your noise floor and your signal's quietest partials. Play the noisy source silent for a moment, raise threshold until the noise just vanishes, then stop — every further dB starts eating signal.

slope — how hard the door closes

The soft knee. 0 passes everything (bypass by another name); low values fade bins gently as they approach the threshold; high values approach a hard per-bin gate. And here lives the genre's famous artifact: push slope hard with threshold high and bins near the boundary flicker open and shut frame by frame — musical noise, a watery, birds-in-the-pipes chirping. The cure is almost always a gentler slope and a lower threshold, accepting a little noise instead of a lot of artifact. Half the craft of spectral gating is knowing when to stop.

FFT size — resolution vs. smearing (and the latency)

The frame size trades three things at once:

  • Frequency resolution: bigger frames separate closely spaced partials from noise between them — better gating for dense, tonal material.
  • Time smearing: bigger frames blur transients; a gate decision spreads across the whole frame. Percussive material wants smaller frames.
  • Latency: exactly one FFT frame, by construction. 2048 samples at 48 kHz is 43 ms — fine on a mix bus, noticeable on a live input.

Recipes

  • Location dialog cleanup: moderate frame, threshold found by the silent-passage method above, slope as low as removes the hiss. Listen to the pauses — that's where both the win and the artifact live.
  • Synth-line de-hiss: tonal material with stable partials is the best case — bigger frames, and the gate closes every bin the notes don't own.
  • Creative abuse: absurd threshold with a hard slope isn't repair, it's an effect — the signal reduced to its loudest spectral bones. The artifact becomes the instrument.

When it is not the right tool

  • Noise under the signal, not beside it. A gate — even per-bin — only removes noise where the signal isn't. Broadband hiss sharing bins with a broadband source needs subtraction/statistical methods, a different machine.
  • Hum and buzz. A 50/60 Hz family is a few known frequencies; surgical notches (tap.filter~) beat a thousand bouncers who all have to guess.
  • Time-domain gating with musical envelope shaping — attack/hold/release on the whole signal is a classic gate's job, and it doesn't smear transients.

Checkpoint

An STFT expander: per-bin thresholds close the spectrum's quiet gaps, the machinery reconstructs bit-faithfully when open (pinned below 10⁻⁶), and the price is one frame of latency plus the musical-noise artifact that appears exactly when threshold and slope are pushed past honest. Find the floor, close the door gently, and stop while the pauses still sound like air instead of water.

The spectrum, re-plumbed

Every process so far in this book treats the spectrum with respect: filters shade it, gates prune it, vocoders dress it up. tap.spectra~ re-plumbs it. Each output bin k is filled from input bin round(k · remap) — the spectrum's contents redistributed by a rule with no acoustic justification whatsoever. It is the one object in this book whose purpose is to sound like nothing in nature, and it is honest about it: the reference page has called it an "ultra-non-linear effect" since 2002. This chapter is what the rule does, why the results are inharmonic almost everywhere, and how to drive an effect whose sweet spots are narrow and strange.

Companion material: the reference page and help patcher in the TapTools-Max package; the kernel's Catch suite pins the two anchor behaviors below.

The rule, and its two pinned anchors

Inside the object's own STFT (the same Hann/4×-overlap engine as tap.nr~), the lower half of the output spectrum is assembled by reading input bins at k · remap, and the upper half is mirrored to keep the spectrum Hermitian — so the output is always real, whatever violence the remap did. Two behaviors are pinned by test:

  • remap 1 is the identity: the output reconstructs the input exactly, delayed by one FFT frame. Transparent machinery, like its sibling.
  • remap 2 moves input bin 2k to output bin k — the spectrum compressed toward the bottom: content from twice the frequency lands at half.

Why almost everything comes out inharmonic

Pitch shifting scales frequencies continuously; this remaps bin indices, quantized to round(k · remap). A harmonic series at f, 2f, 3f… survives integer remaps in recognizable form — remap 2 folds a harmonic spectrum roughly an octave down — but at remap 1.37 the partials land on a grid nature never drew: some merge, some vanish, spacings go irrational-ish. The result reads as bells, metal, ghosts of the input. That in-between space is the instrument. Sweep remap slowly across 1.0 and you can hear the sound leave reality and come back.

Two practical corollaries:

  • remap just above or below 1 (0.9–1.1) is the subtle zone — a detuned, phasey shadow of the input, cheaper than it sounds.
  • remap well below 1 stretches the low spectrum upward across the output (each output bin reads a lower input bin), thinning the top; well above 1 compresses everything into the bass and discards the input's top octaves entirely. Loud, dark, and blunt — usually wants a fresh brightness source afterwards.

The knobs

There is really one, plus the frame:

remap — the rule

Continuous. Identity at 1; integer values are the quasi-musical landmarks; everything between is the inharmonic wilderness. Automate it slowly — the per-frame quantization means fast sweeps step audibly, which is either the problem or the point.

FFT size — the grain of the grid

Bigger frames put the bins closer together, so the remap grid is finer: less quantization grit, smoother inharmonicity, more latency (one frame, as always) and more transient smearing. Smaller frames make the remap chunkier and more overtly digital. Unlike tap.nr~, where the frame is a fidelity question, here it is a flavor question.

Recipes

  • Bell foundry: harmonic material (a tap.vco~ saw, a piano) at remap 1.3–1.6, into a long reverb. Instant inharmonic percussion.
  • The shadow voice: speech at remap 0.95, mixed subtly under the dry — a wrongness the ear notices before the mind does.
  • The corridor: automate remap 1.0 → 2.0 over a minute under a sustained chord — a slow departure from consonance that lands, at exactly 2, somewhere almost stable again.
  • Stacked plumbing: two in series at remap a and b is a remap at a·b with two layers of quantization grit — the grit is the reason to do it.

When it is not the right tool

  • Musical transposition. The remap is spectral plumbing, not pitch shifting: use tap.shift~ for clean intervals, tap.pitchaccum~ for the spiral.
  • Harmonizing or formant work. Nothing here knows what a formant is; the rule moves bins, not vowels.
  • Subtle timbre correction. Even at its gentlest this object is a character effect; EQ-shaped intentions belong with tap.filter~ or tap.svf~'s EQ modes.

Checkpoint

One rule — output bin k reads input bin round(k · remap), Hermitian-mirrored — inside a transparent STFT: identity at 1 (pinned), octave-fold at 2 (pinned), and an inharmonic wilderness everywhere between. The FFT size sets the grain of the grid, the sweet spots are narrow, and that is the appeal: this is the book's one unapologetic reality-distortion tool. Use it where nature's spectra have gotten boring.

The acid machine

The Roland TB-303 was designed to imitate a bass guitar, failed completely, and accidentally defined thirty years of dance music. What makes it unmistakable is not any one block — a saw into a lowpass is every synth ever made — but the coupling: accent drives the filter and the amplifier through shared circuitry with memory across notes, slide is a gate that refuses to let go, and the envelopes are fixed RC discharge curves with exactly one knob between them. tap.303~ is a circuit-informed model of that whole tangle; tap.diode~ is its filter as a standalone object; tap.303.seq~ is the other half of the instrument. This chapter is what each control trades, and what the measurements say the model actually delivers.

Companion material: the reference pages and help patchers in the TapTools-Max package, and two executed verification notebooks — tb303.ipynb for the voice and step_seq.ipynb for the sequencer — every number below is a measurement from one of them or from the kernel test suite. Provenance runs through Tim Stinchcombe's filter analysis, Robin Whittle's Devil Fish documentation, the x0xb0x schematics, and Robin Schmidt's Open303, whose measured calibrations several constants adopt verbatim.

What the hardware is, in one paragraph

One saw-core oscillator (the "square" is the saw through a transistor shaper, not a clean pulse), into a four-stage diode-ladder filter — not the Moog transistor ladder; the diode ladder's stages load each other, which is why its resonance is broader, less pure, and entirely its own — then a one-transistor amplifier. Two envelopes, both decay-only RC discharges: the Main Envelope sweeps the cutoff (the envmod knob decides how much), the VCA envelope is fixed. Accent makes the Main Envelope hotter and faster, routes it into the VCA, and charges a capacitor (C13) through the resonance pot — and because C13 doesn't fully discharge between closely spaced accents, runs of accented notes bloom, the famous wow. Slide holds the gate across the step boundary while the pitch CV glides through a ~60 ms RC. Everything about a note — pitch, gate, accent, slide — comes from the sequencer, not the panel. That is why this is three objects, not one.

The filter first: tap.diode~

The panel says "18 dB/oct"; the circuit is four poles whose asymptotic slope is 24 dB/oct with a shallower region near cutoff — Stinchcombe untangled this, and the kernel reproduces his published transfer function to 0.028 dB. Two behaviors are load-bearing and easy to get wrong:

  • The resonance feedback runs through a 150 Hz high-pass, so resonance thins as the cutoff drops — low notes squelch, they don't ring. Pinned by test: the ring-down Q falls with cutoff.
  • A stock 303 never quite self-oscillates, and neither does this filter at stock settings. That emerged from the modeled feedback high-pass rather than being programmed in, and it's documented as a trait, not a defect. (Push resonance past 1.0 — the bend range runs to 1.5 — and it will sing for you anyway.)

Like tap.ladder~ it has a solver choice: fast (default) or exact, which iterates the re-linearized solve to convergence on the true nonlinear loop. Measured across a matrix out to resonance 1.4 and +24 dB drive — beyond anything the hardware can reach — the two differ by at most −44.9 dBr, at 1.6–3.3× the CPU. The exact solver is there for the suspicious; the fast one is there for the patch. oversample (1/2/4, default 2) and a signal-rate cutoff in the right inlet round out the tap.ladder~ surface.

The voice: tap.303~, knob by knob

The attributes mirror the seven-knob panel; the calibrations are Open303's measured laws.

Signal-flow diagram of the 303 voice: pitch through slide into oscillator, shaper, coupling highpass, diode ladder, and VCA, with the accent bus fanning to the envelope, the C13 sweep capacitor, and the VCA, and the envelope-driven cutoff CV feeding the ladder

The blocks are ordinary; the red and amber wires are the 303. Accent touches three destinations at once, and C13 remembers across notes — the couplings are the instrument.

  • waveform — saw or square. The square is the hardware's shaped saw: −tanh(10^(36.9/20)·saw + 4.37), Open303's measured constants verbatim — rounded and notched, audibly not a 50 % pulse.
  • cutoff — the knob in Hz. Stock travel is the measured 302–2394 Hz; the attribute range (100–5000) is a flagged bend beyond the panel.
  • resonance — 0..1 is stock; up to 1.5 is the bend.
  • envmod — how much Main Envelope reaches the cutoff, with the hardware's measured law: 2/3 of the sweep goes above the knob position, 1/3 below, and the "gimmick" offset shifts the resting point down as you turn it up. The knobs feel right because the interaction is modeled, not just the ranges.
  • decay — Main Envelope decay, 200 ms–2 s. On an accented note the hardware ignores this knob and runs at ~200 ms; so does the model (adjustable via the accdecay bend, 50–2000 ms).
  • accent — how hard accented notes hit: louder and punchier (the envelope routing), and quackier (the C13 sweep, scaled by the resonance knob). The wow is measured: over a run of closely spaced accents the cutoff peak builds by ×1.94, and decays back within ×0.998 once the accents stop. Consecutive accents at high resonance are the entire genre.
  • tuning, gain — cents and dB. Plumbing.

The envelopes carry the schematic's fixed interrelations: MEG attack ~3 ms, VCA attack ~3 ms with a measured ~1.23 s decay chopped at gate-off, 50 ms when accented. None of these have knobs on the hardware, so none of them have knobs here — except through the documented Devil-Fish-style bends (slide 10–500 ms, attack 0.3–30 ms, accdecay, and drive ±24 dB into the ladder, where the diodes compress: +24 dB of gain buys only 9.2× of RMS). All stock at their defaults.

Phase 2 added vca clean|warm: the one-transistor class-A stage as a slope-normalized biased saturator, in the hardware's signal order. The distortion tracks the envelope — measured 5.4 % difference signal on quiet notes, 11.5 % on hot accents — so warm thickens exactly where the hardware does. clean (default) is bit-identical to phase 1.

House machinery throughout: seed/tolerance per-unit component spread (an mc. stack of 303s with different seeds detunes and drifts like a wall of real units), 16 preset-morph slots with factory acid in 1–8 (squelch, sub, screamer, rubber, knock, bloom, overdriven, glass), and per-sample ramps on every parameter.

The note interface, and why slide is free

tap.303~ is TapTools' first pitched instrument, and its inlets are the package-wide melodic contract: pitch as a MIDI note number signal in the left inlet, gate with amplitude-as-accent in the right — 1.0 is a plain note, 2.0 fully accented (depth = amplitude − 1). Slide needs no input at all: a pitch change while the gate is held is a slide — legato, no envelope retrigger, the ~60 ms RC glide — which is exactly the hardware's own definition. A note <pitch> [accent] [slide] message covers patching without signals.

The other half: tap.303.seq~

Half the 303's sound is sequencer behavior, so the sequencer emits the voice's contract verbatim: a pitch signal and a gate signal, clocked by a phase ramp (0..1 per pattern, a phasor~). Per step: pitch, gate/rest, accent, slide. The measured facts, from the sequencer notebook:

  • Steps land on the analytic grid within one sample; the gate opens at the step start and closes at 0.5 of the step (Open303's stepLength).
  • A slid step is approached with the gate held: 16 gated steps with 3 slide flags produce exactly 13 note-ons — the other three arrive legato, pitch stepping on the boundary sample, and the voice glides.
  • Accented steps gate at 2.0; transpose shifts live, like the hardware's transpose mode without the mode; swing and pattern slots with cycle-quantized recall are shared with the drum rows (next chapter).

The 1981 pitch-mode/time-mode data entry is deliberately not recreated. You keep the data model; you lose the part everyone hated.

When it is not the right tool

  • You want a generic bass synth. tap.vco~ + tap.svf~ + tap.adsr~ give you ADSRs, waveform variety, and a filter that behaves. This object's value is its refusal to decouple.
  • You want the filter without the biography — tap.diode~ alone, or tap.ladder~ if you want the Moog character instead of the 303's.
  • You want polyphony. It's a monosynth; mc. gives you many monosynths, which is not the same thing as a polysynth and shouldn't be.

Checkpoint

A diode ladder that matches the published analysis to 0.028 dB and won't self-oscillate until you bend it; a voice whose envelopes, accent path, and C13 memory come from the schematic, with the wow measured at ×1.94 across an accent run; slide as pure gate-hold, so legato falls out of the note contract; and a sequencer that emits that contract sample-accurately. The coupling is the instrument — and every claim above has an executed notebook cell behind it.

The drum machine

The Roland TR-808 is the most thoroughly analyzed drum machine in the academic literature, and the reason is charming: the whole instrument is analog synthesis. No samples anywhere — every sound is a small circuit, and most of them are variations on about four ideas. The tap.808.* family recreates the eight voice channels circuit block by circuit block, one external per channel, and tap.808.seq~ supplies the machine's other half as one sequencer row per patch cord. This chapter is the family tour: the shared trigger contract, each voice's character and knobs, and the calibration pass against a real unit that the numbers come from.

Companion material: each voice's reference page and help patcher, the family overview patcher (tap.808.maxhelp — all eight voices sequenced off one phasor~), the tr808_calibration.ipynb notebook, and the step_seq.ipynb sequencer notebook. Provenance runs through the Werner–Abel–Smith papers (DAFx-14 and companions) and the TR-808 Service Notes, read component by component; every magic constant in the kernel headers carries its schematic designator.

One trigger to rule them all

On the hardware, every voice hangs off a common trigger bus: the CPU's 1 ms pulse rides a voltage between 4 and 14 V depending on the accent circuit, and a hotter pulse excites each circuit harder — more punch, slightly different timbre — not merely louder. The family keeps that literally: every voice fires on a signal rising edge, and the edge's amplitude (0..1) is the accent, mapped onto the 4–14 V bus. bang and trigger 0.7 messages cover the scheduler side. Filter states persist across triggers, so fast rolls interfere with the ringing tail like the hardware — no machine-gun effect. And because the excitation is a voltage, anything that makes an edge can play the kit: a click~, an envelope, a tap.303.seq~ gate, or the row object built for the job.

The voices

tap.808.kick~ — the bridged-T with a biography

The bass drum is a damped bridged-T resonator (~49.4 Hz from the modeled component values; Roland's chart optimistically says 56, real units measure as low as 48) with three behaviors that make it the kick, all emergent from the modeled schematic: for the first ~6 ms the envelope saturates Q43 and the resonator sits near ~129 Hz — the attack punch, which is a different mechanism from the famous downward pitch "sigh" (leakage through R161, the paper's fitted nonlinearity); and a retriggering pulse re-excites the center node as the envelope collapses so the note doesn't step down. Panel knobs: decay (seconds of ring at the top), tone (click at ~7 kHz down to ~300 Hz), level. Paper-documented bends, stock by default: tuning, pulse, sigh, attack — turn sigh 0 and the pitch relaxation disconnects, exactly as the bend does on the bench.

Calibration: against a real unit's knob-gridded sample set, the fundamental sat within 2.4 % at every tone/decay position and the −40 dB decay endpoints within 6 % (72 ms → 2.36 s measured, 69 ms → 2.42 s modeled) — no constant needed changing.

tap.808.snare~ and tap.808.clap~ — resonators plus noise

The snare is two bridged-Ts (the late-revision ~173/336 Hz pair) with a trigger divider and the "snappy" path — enveloped noise, band-limited around 4 kHz to the measured unit. Fundamentals calibrated within 1.2 %, including the mode flip at tone-max. The clap channel (@model clap|maracas) is the Service Notes' Figure-13 circuit: band-passed noise near 2 kHz through a VCA driven by a three-teeth sawtooth retrigger — the "multiple hands" transient — plus the Q70 reverberation tail. The maracas mode is the same noise voiced short and bright.

tap.808.hat~ and tap.808.cymbal~ — the metal bank

Six Schmitt-trigger square oscillators (205.3, 369.6, 304.4, 522.7 Hz plus the two trimmer-tuned at 800 and 540, duty 47.98 %) feed two bandpass voicings near 3.4 and 7.1 kHz. Werner et al. measured that resistor variance puts any given unit up to ~20 % off those frequencies — which is why no two 808s' cymbals sound alike, and why seed/tolerance exists: every seed is a different unit off the line, and an mc. stack of cymbals decorrelates like real hardware. The hats are one object with two trigger inlets because on hardware they are one circuit with two envelope paths and a choke — closed chokes open (the Q23/R173 path), pinned by test, and unimplementable as separate externals. Open-hat decay spans the chart's 90–600 ms; the cymbal's two separately enveloped bands cover its 350–1200 ms "sizzle" span. tap.808.cowbell~ taps just the 540/800 pair into the ~860 Hz voicing with a two-slope envelope; more cowbell is a patching decision.

tap.808.tom~ and tap.808.rim~ — the resonator variations

Six sounds on two objects, as the hardware switches them: @size low|mid|high × @model tom|conga. Congas are the tom circuit without its noise layer, tuned differently; the toms add the D80/D81 attack pitch fall and a pink noise layer. The rim channel is @model rimshot|claves: the rimshot's ~1667 + 455 Hz crack with the swing-VCA's harmonics, versus the claves' pure ~2500 Hz tick. Tunings sit within ~4 % of the measured unit.

The bridged-T resonator circuit: an op-amp with capacitive arms, a resistive bridge, and a leg to ground, triggered through an injection resistor, with the kick's per-sample leg modulation drawn in red — and the eight voices grouped by how they use it

Roland's universal voice circuit. Eight voices, one network — the kick earns its punch by modulating the leg per sample.

The calibration pass, honestly

The §7.2 calibration ran against a real TR-808 (s/n 103852) recorded from the individual outs with knob positions encoded in the filenames — a 0/2.5/5/7.5/10 dial grid, 116 samples — which upgraded "sounds right" to a quantitative per-knob-cell comparison. Identical measurements (spectral-peak fundamental, −40 dB decay, power centroid) ran on both sides. The pitches were already right nearly everywhere; what the pass actually changed was time: tom, conga, cowbell, and clap tails roughly doubled to match the unit, the snappy was band-limited and re-enveloped, the rimshot re-voiced low-dominant, the cymbal's decay span corrected. Each kernel header carries its residuals. The lesson generalizes: schematics get you the frequencies; recordings get you the envelopes.

The other half: tap.808.seq~

One row of the 16-step sequencer, as an object: feed it a phase ramp (0..1 per pattern, a phasor~) and it emits trigger impulses whose amplitude is the step's accent — the family contract, straight into any voice. Twelve rows off one phasor are the hardware's panel, sample-locked forever; the accent row falls out of giving every row the same accents list. The measured facts, from the sequencer notebook: steps land on the analytic grid within one sample; the pinned levels are plain 0.01 (the 4 V base — an un-accented hit still strikes the circuit) and accented 0.5 (the accent knob at noon; 1.0 is the full 14 V); swing delays the off-16ths by exactly swing/2 of a step; a length 12 row against 16s is the triplet pre-scale generalized to polymeter; and pattern slots with cycle-quantized recall are the A/B-half and fill switching as one message. pulse widens the impulse into a held gate when you'd rather drive tap.adsr~ than a drum.

When it is not the right tool

  • You want a kick, not the kick. A sine with an envelope is cheaper and takes EQ more politely. This family's value is the circuit behavior — the attack jump, the choke, the accent-as-voltage.
  • You want your own drum sounds. These circuits are what they are; seed, the documented bends, and the panel knobs bend them, but a sampler is a sampler.
  • You want 909 hats. The 909's metal is sampled; this machine's is six square waves. Different instrument, different chapter, maybe someday.

Checkpoint

Eight channels, four circuit ideas — bridged-T resonators, a shared metal bank, noise paths, swing-VCAs — under one amplitude-as-accent trigger bus, calibrated per knob cell against a real unit and honest about what changed (the tails) and what didn't (the tunings). The hats choke because they share a circuit; the cymbals decorrelate because resistors do; and the sequencer row emits the same voltage idea the voices drink, so the whole kit runs off one phasor ramp. The machine's two halves, both measured.

The note you meant

Every sung note is two notes: the one that happened and the one you meant. tap.tune~ measures the distance between them and closes it — how fast it closes it is the whole instrument. Closed slowly, nobody knows it was there. Closed instantly, everybody knows: that snap is the most famous vocal effect of the last twenty-five years. One object, one time constant, both worlds.

A short history matters here, told plainly. The classic pipeline — detect the pitch, snap it to the nearest allowed note, retune by time-domain resynthesis — was patented in 1998 and the patent expired in 2018, which is why a whole field of tuners exists today and why this object can implement the technique from the literature. The famous product name remains a live trademark, which is why this object is called tap.tune~ and this chapter says "hard snap" instead. And editing individual notes inside a chord remains patent-fenced territory — tap.tune~ is monophonic by design, not by omission. (None of this paragraph is legal advice; the project's own ship-gate is a freedom-to-operate review.)

Companion material: the reference page and help patcher in the TapTools-Max package, the runtime maxtest, and two executed notebooks — tune.ipynb here and pitchshift.ipynb in the DspTap repo — that measured every claim below.

Signal-flow diagram of tap.tune~: a per-hop brain of YIN, target mapper, glide, and ratio over a per-sample path through the input ring into the selectable resynthesis backend

A per-hop brain over a per-sample corrector, with one seam where three resynthesis engines interchange.

The knob that is the instrument: speed

speed is the time constant, in milliseconds, of the glide onto the target note.

  • 0 ms — the hard snap. The correction lands within a detection hop (~5 ms). Vibrato gets quantized into terraces; note transitions become instant staircase steps. This is the effect, worn on the outside.
  • 10–40 ms — classic correction. Fast enough that a listener hears "a singer with good intonation," slow enough that the attack of each note — where identity lives — is not robotic. The default is 20.
  • 100 ms and up — intonation leaning. The corrector arrives so late it only tames drift; vibrato passes through nearly untouched.

The notebook's pitch-track figure shows all three glides onto the same 46-cent-sharp note; the kernel test pins the exponential's arrival. There is also amount (0–100%): a fader on the correction distance itself. 100 lands on the target; 50 splits the difference — a gentler kind of honesty that keeps a performance's shape while shrinking its errors.

Telling it what is allowed

The corrector never invents a target; it snaps to the nearest note you allowed.

  • key + scale — the usual contract: @key d @scale major and every detected pitch pulls toward the nearest D-major degree. Presets: chromatic, major, minor, harmonic, melodic, pentatonic, minorpentatonic.
  • notes — the twelve toggles, absolute pitch classes C through B, panel-style: notes 1 0 0 0 1 0 0 1 0 0 0 0 snaps everything to a C-major triad, which is less a correction than an arrangement decision.
  • mode midi — the target is the nearest currently held MIDI note (note 64 100 holds E4; velocity 0 releases; flush clears). Hold one note and everything becomes that note; hold a changing chord's roots and the corrector is suddenly a performable melody-mangler. No notes held means no correction — the object never guesses.

An empty mask behaves the same way: nothing allowed, nothing changed.

Three engines, one corrector: backend

Detection, targeting, and the glide are shared; only the resynthesis swaps. All three land the same intonation — the notebook drives the same vibrato "voice" through each and all three settle on 220.00 Hz — so the choice is about character and latency, not accuracy.

backendwhat it ischoose it forlatency @ 48 kHz
graintwo-tap delay-line, window locked to the detected period (the tap.shift~ engine)the default; lowest latency, waveform-preserving, happy on any materiala few ms
psolatrue TD-PSOLAvoice — it preserves formants by construction~36 ms
pvocpeak-locked phase vocoderdense, harmonically rich material; pairs with formant~21 ms

Switching live is click-safe: the incoming engine starts from silence and fades in rather than splicing stale audio. One honest caveat per engine: grain colors sustained unpitched input with a mild moving comb (the known trade of its class); psola wants harmonic material — on a pure sine shifted far, its output legitimately thins (the machine chapter explains why that is the same property as its formant preservation); pvoc smears sharp transients slightly, as every phase vocoder does.

Keeping the singer's mouth: formant

A correction of thirty cents moves formants thirty cents — nobody hears it. A MIDI-mode command of five semitones moves them five semitones — everybody hears it; that is the chipmunk. @formant 1 enables LPC formant preservation on the pvoc backend: the pitch moves, the vocal tract's envelope stays where the singer put it. The notebook corrects a synthetic voice up 5.5 semitones both ways; with the flag on, the formant bump stays put (band-energy ratio 730:1 in its favor). psola needs no flag — formant preservation is its resampling rule — and grain ignores the flag.

Letting it find the key: autokey

@autokey 1 starts a learner: every voiced detection drops its pitch class into a histogram that forgets with about a minute of memory, scored against the published Krumhansl–Kessler key profiles. Two design decisions worth knowing:

  • It never acts on its own. A key estimate that silently re-aimed your targets mid-phrase would be a bug wearing a feature's clothes. getkey asks (the right outlet answers key d major 0.95, or key none in the first half-second); applykey adopts the estimate into the key and scale attributes — visibly, where you can see and undo it.
  • It forgets on purpose. The one-minute memory means a modulation stops arguing with the old verse about as fast as you stop playing it.

The kernel test plays a D-major scale and reads back D major at 0.95 confidence; an A harmonic-minor melody reads as A minor.

The right outlet

While the input is voiced, the right outlet reports pitch <midi> <hz> every @interval milliseconds (default 50; 0 disables; a pitch -1 0 marks the end of voicing). That is a free tuner display, a melody recorder, or the control signal for whatever you want to drive with the singer's pitch — and it is the same detector the corrector itself uses, so what you see is what it acted on.

Recipes

  • Invisible repair: @scale major @key (your key) @speed 25 @amount 80. The 80 keeps a little humanity in the intonation; nobody will name what changed.
  • The famous one: @speed 0 @scale minorpentatonic. Fewer allowed notes make the terraces wider and the snap prouder. Add melisma.
  • One-note choir: @mode midi @speed 5, hold a note, feed it speech. Everything becomes chant on that pitch.
  • Formant-true transposer: @mode midi @backend pvoc @formant 1 @speed 10, play a melody against a held vocal — a harmonizer that keeps the singer's identity.
  • Tuner display only: @amount 0 @interval 20 — the object corrects nothing and the right outlet becomes a clean pitch stream.

When it is not the right tool

  • Chords. The detector is monophonic; a chord reads as garbage or as its loudest note, and per-note polyphonic editing is deliberately out of scope (see the history paragraph). Split voices first, or don't.
  • Drums, breath, speech consonants. Unpitched input passes through with no correction — by design — but the grain engine adds its mild comb coloration to sustained noise. For processing unpitched material there are better rooms in this house.
  • Creative shifting. If the goal is an interval rather than intonation, tap.shift~ is the plain shifter, tap.harmony~ the formant-preserving chord stack, and tap.pitchaccum~ the spiral; tap.tune~ always measures first and that measurement is latency you don't need.

Checkpoint

Detect, snap to the nearest allowed note, glide at speed — that is the whole machine, and speed is the dial between honesty and effect. Targets come from key + scale, twelve toggles, or held MIDI notes; three resynthesis engines trade character against latency while landing the same intonation; formant keeps the singer's mouth in place when corrections get big; autokey learns the key but only ever suggests. The right outlet tells you what it heard. And when the input isn't a single pitched voice, the honest move — which the object makes — is to change nothing.

Distortion with a memory

Every distortion plugin can bend a transfer curve. tap.overdrive~ is built on the observation that the pedals people actually love — the Tube Screamer lineage, and specifically the Mad Professor Little Green Wonder that served as this object's listening reference — don't apply one curve to the whole spectrum. Their clipper lives inside an op-amp's feedback loop with frequency-dependent parts around it, and that loop is most of the sound: bass sees less gain and stays tight, mids break up first, and the knee never quite flattens because the clean signal always rides through. A memoryless waveshaper — including both modes of the Jamoma-era tap.overdrive~ this object succeeds — structurally cannot do any of that. This one can, because the shaper sits inside a lowpass feedback loop: distortion with a memory.

Companion material: the reference page and help patcher in the TapTools-Max package, and the verification notebook, where every number below is an executed, plotted measurement of the shipping kernel. The figures in this chapter are measurements too — regenerated from the same kernel through the C ABI by book/figures/overdrive.py, never drawn by hand.

What the loop buys

The claim worth leading with, because no static curve can make it: the object's small-signal gain tilts with frequency, and the tilt grows with drive. Measured between 80 Hz and 4 kHz, the tilt is +5 dB at drive 0 (just the voicing EQ), +16.3 dB at drive 0.5, +17.2 dB at drive 0.9. Low frequencies are pinned near-clean by the feedback while mids and highs take the full drive gain — so a low E stays articulate under the same setting that saturates the pick attack. That is the Tube Screamer "tightness" in one plot:

Small-signal gain versus frequency at drive 0, 0.5, and 0.9: the curves tilt progressively steeper as drive rises, with bass pinned and mids lifting

The measured headline. A memoryless shaper's version of this figure is three horizontal lines.

The second structural trait: the transfer never flattens. A unity clean path is summed around the clipper — the non-inverting op-amp topology — so however hard the shaped part saturates, output keeps rising with input (measured strictly monotonic at every drive setting). The old sine-shaper mode's hard ±1 plateau, a large part of what read as "digital," is gone by construction.

Output peak versus input peak at three drive settings: every curve keeps rising with reduced slope, none goes flat

Compression without a ceiling: the slope falls as drive rises, but never to zero.

The knobs, one by one

drive — 0 to 1, edge-of-breakup to saturated

Normalized, like every musical parameter on this object, with the perceptual mapping done inside (the knob sweeps the clipper's gain from +6 to +46 dB, with a level compensation tracking it). drive 0 is a pedal's gain knob at full counterclockwise — still warm, not bit-clean; bypass is the clean switch. The normalized range maps directly onto MIDI/OSC controllers, and onto Q15/Q31 fixed-point for the embedded ports this kernel is written to survive.

body — the signature voicing control

The LGW's defining knob, reproduced as linear pre/post EQ around the clipper (that's what it is in the pedal — voicing, not nonlinearity). Toward −1, fuller lows reach the clipper and the top gets a slight shelf lift; toward +1, the lows thin and tighten and an upper-mid bell pushes forward — centered at 1150 Hz, deliberately above the classic TS hump. Measured at the extremes: 100 Hz moves by 10 dB, the 1150 Hz push adds 4 dB, the counterclockwise treble lift is +2.5 dB at 8 kHz. The exact centers and gains are by-ear placeholders pending the in-Max voicing pass against LGW demos — the shape of the control is final, the seasoning isn't.

Small-signal response at body −1, 0, and +1: fuller lows and a top lift counterclockwise, thinner lows and an upper-mid push clockwise

The knob's whole range. Note the crossover around 500 Hz: body trades lows against upper mids around a stable center, like the pedal.

asymmetry — the even harmonics the old object couldn't make

Both Jamoma modes were odd functions: odd harmonics only, the entire "warmth" vocabulary absent. asymmetry biases the clipper: at 0 the path is exactly symmetric (measured H2 at −151 dB — the numerical floor), and raising it brings the even series up smoothly (H2 at −26 dB by asymmetry 0.6). The default sits at 0.15, a small nonzero warmth chosen by ear. Asymmetric clipping generates DC, so a DC blocker sits permanently after the clipper — measured output mean under full drive, full asymmetry: 10⁻¹⁰. (The original TTOverdrive contained a DC blocker whose output was computed and then discarded; this one is load-bearing.)

Harmonic spectra of a 220.5 Hz tone at asymmetry 0 and 0.6: the left panel shows odd harmonics only, the right adds the full even series

The same tone, the same drive — the only change is asymmetry, and the even series (H2, H4, …) appears between the odd lines.

oversample — 1, 2, 4, or 8; default 4

Clipping makes harmonics; harmonics past Nyquist fold back as inharmonic junk. At 1× a hard-driven 5 kHz tone puts its folded seventh harmonic at −22 dB relative to the fundamental — clearly audible garbage at 12993 Hz. At the default 4× the same component measures −36 dB, with the true harmonics unchanged. Turn it down to 1× only when CPU matters more than the top octave, or when you want the fizz.

preamp, output, smooth, bypass, mute

Input and makeup gain in dB (±24) — the only unit-bearing parameters, because gains are the one place real units belong. Everything ramps click-free over smooth milliseconds (default 20).

Where it sits in a patch

Mono by design; wrap it in mc. for multichannel like the rest of the package. It takes line-level signals as happily as guitar DI — the drive mapping is normalized to full-scale digital, not to pickup output. For the LGW move, start at drive 0.4, body -0.3, asymmetry 0.15 and ride body against the source's low end. For a clean boost that just thickens, drive 0 with asymmetry 0.3. For fuzz territory this is the wrong object on purpose — the loop keeps pulling it back toward articulation.

Every claim above is pinned twice: as an executed measurement in the notebook, and as a hard assertion in the kernel's Catch2 suite (tests/overdrive_test.cpp), which CI runs on every push. The math behind the loop — including why it had to be solved zero-delay, and what happens if you don't — is in the machine chapter: The clipper in the loop.

Solving the filter on paper: svf.h

The user-facing chapter promised that tap.svf~ is "unconditionally stable under per-sample cutoff modulation" and that its morph corners are "bit-identical to the discrete modes." Promises like that are either mathematical facts or marketing. This appendix does the math: it derives the filter the way the file was actually designed, then walks the engineering decisions that don't show up in a Bode plot — and why each one beat its alternative.

The reference is Andy Simper's Cytomic technical papers ("Solving the continuous SVF equations using trapezoidal integration and equivalent currents"), specifically the SvfLinearTrapOptimised2 form. What follows is the same derivation with the file's variable names.

The analog prototype, and where digital versions go wrong

The state-variable filter is two integrators in a loop. With cutoff ω and damping k = 1/Q, the continuous equations are:

v1' = ω · (v0 − k·v1 − v2)     (band state: input minus damping minus low state)
v2' = ω · v1                   (low state: integral of band)

Lowpass is v2, bandpass v1, highpass v0 − k·v1 − v2 — every response lives in the same two states, which is what makes an output mix (and therefore a morph) possible at all.

The textbook digital version (Chamberlin) discretizes with explicit Euler: each integrator uses the previous sample's value. That inserts a unit delay into the loop, and a delay in a feedback loop is a stability bomb with a frequency fuse: the design blows up as fc approaches fs/6, and modulating the cutoff re-lights the fuse every sample. The classic workarounds (oversample it, clamp it) treat symptoms.

Trapezoidal integration, and the algebraic loop

The fix is to integrate with the trapezoidal rule — average the old and new derivative — which in filter terms is the bilinear transform. Define the prewarped gain the file computes once per cutoff change:

g = tan(π · fc / fs)

(The tan is the prewarp: it makes the digital filter's response at fc exactly match the analog prototype's, all the way to Nyquist. This single line is why self-oscillation later measures 999.7 Hz for a 1 kHz setting rather than drifting flat.)

Trapezoidal integration of state s with input x is s_new = s + g·(x_old + x_new). Grouping the "old" terms into a memory variable — Simper's equivalent current ic = s + g·x_old — each integrator becomes:

s_new = ic + g · x_new         with ic updated as ic_new = 2·s_new − ic

But notice the trap: x_new for the first integrator is the new band value, which depends on the new low value, which depends on the new band value. The new sample appears on both sides — an algebraic loop, exactly the "zero-delay feedback" the initials ZDF refer to. Instead of breaking the loop with a delay (Chamberlin's sin), we solve it. It is linear, so substitution gives a closed form. With v0 the input and ic1, ic2 the two equivalent currents:

v1 = a1 · ic1 + a2 · (v0 − ic2)          the band state, solved
v2 = ic2 + g · v1                        the low state, then follows
where  a1 = 1 / (1 + g·(g + k)),  a2 = g · a1

Those are precisely the file's per-section solve constants (a1, a2, a3 = g·a2 caches the product used by the low state), and the state update is the canonical TPT pair ic1 = 2·v1 − ic1, ic2 = 2·v2 − ic2.

Why this is unconditionally stable, even modulated: the trapezoidal rule is A-stable — it maps the entire left half of the s-plane (every stable analog filter) inside the unit circle, for any g > 0. Change g every sample and each sample still computes a passive, energy-consistent step; there is no regime of fc or modulation rate where the update gains exceed unity. The notebook's 90 Hz-LFO-through-five-octaves torture test isn't surviving by margin; it's surviving by theorem.

The TPT SVF core as a diagram: two trapezoidal integrators, the damping and low feedback into the input sum, and the downstream output mixer

The loop the algebra just solved, and the mixer the next section explains.

The output mix, and why morph corners cost nothing

Every response is a weighted sum over the same solved values:

y = m0·v0 + m1·v1 + m2·v2

lowpass   m = (0,  0, 1)        notch     m = (1, −k, 0)
bandpass  m = (0,  1, 0)        peak      m = (1, −k, −2)
highpass  m = (1, −k, −1)       allpass   m = (1, −2k, 0)

mode_morph linearly interpolates the mix vector around the circle LP → BP → HP → notch → LP. Two facts follow by construction, not by tuning:

  • At a corner, the interpolated vector equals the discrete mode's vector exactly — same floats, same states, same arithmetic. The notebook's measured max difference of 0 is not a tight tolerance; it is an identity.
  • Morphing is free. The states don't know the mix exists; sweeping it can never destabilize anything, because it is three multiplies downstream of the filter.

The parametric-EQ trio (bell, shelves) is the same machinery with mix weights that depend on a gain factor A = 10^(dB/40), straight from Simper's tables — and it always runs a single section, because cascading an EQ stage squares its boost: two +12 dB bells are a +24 dB bell, which is never what the user typed.

The cascade: Butterworth spread, resonance on the last section

Orders 4 and 8 run two and four sections at the same cutoff. Stacking identical Q = 0.707 sections would droop the passband (each contributes its −3 dB early); instead the sections take the Butterworth Q spread — the Qs whose product of section responses is maximally flat:

Q_i = 1 / (2·cos θ_i),  θ_i the Butterworth pole angles
order 4: 0.5412, 1.3066        order 8: 0.5098, 0.6013, 0.9000, 2.5629

That is why the measured response sits at −3.01 dB at fc at every order. User resonance then sharpens only the final (highest-Q) section, via

Q_res = Q_base / (1 − r),  r ∈ [0, 1)   (clamped at 1 − 10⁻⁴)

— one clean resonant peak riding a flat passband, rather than four peaks compounding. The inverse mapping (resonance_from_q) exists so the wrapper's q message round-trips exactly.

The driven circuit: one saturation, one pass

The driven circuit places tanh on the band node — in the damping path, where an OTA's transconductance actually compresses. That placement is the whole design: as amplitude grows, the effective damping k·tanh(v1)/v1 grows with it, which is an automatic gain control wrapped around the resonance. Push the loop gain slightly past the oscillation threshold at resonance 1.0 and the filter must oscillate (the linear model's poles are outside the circle) but cannot run away (the saturation restores effective damping as amplitude rises). Bounded self-oscillation is not a limiter bolted on; it is the fixed point of that tug-of-war. An all-zero state solves the equations too — hence "give it a ping."

Solving a nonlinear zero-delay loop exactly needs iteration. The file uses the one-pass scheme shared with tap.ladder~'s solver_fast: solve the linear ZDF prediction for the band node, saturate it, commit. The error of that shortcut is second-order in how much tanh bends over one oversampled step — and the driven circuit always runs oversampled (2× default), with 4th-order Butterworth anti-image/anti-alias biquad pairs on the way up and down. At these rates the one-pass and iterated answers are audibly identical; the ladder file, which drives its nonlinearity much harder, is the one that also ships a Newton option.

The engineering ledger

Decisions visible only in the code, with their reasons:

  • Two-tier coefficient update. Recomputing everything per sample costs a tan() plus the mix logic even when nothing changed. The file splits state: a shape tier (damping, mix weights, EQ gains — dirtied only when a non-frequency parameter or mode changes) and a cutoff tier (tan and the three solve constants — recomputed only when the incoming cutoff differs from the cached one). Signal-rate modulation pays for exactly what it moves. The benchmark ratchet recorded the win: modulated 2nd-order lowpass 36 → 19 ns/sample, modulated morph 77 → 28, bit-identical output (the morph-corner identity tests pin that "bit-identical" is literal).
  • ramp_to doesn't dirty the shape tier for frequency — the cutoff cache catches it. One branch, measurable at audio rates.
  • Multichannel by frame protocol. Coefficients are computed once per tick() and shared by every channel's process(ch, x) — an N-channel engine outside Max for the cost of one solve. The Max wrapper stays mono by house rule (mc. wraps it).
  • Allocation discipline. The only allocation is the per-channel state vector in prepare(); setters are wait-free and safe from the message thread while audio runs, because a "set" is a ramp target plus a dirty flag.
  • Anti-denormal guard on the states (the tap.comb~ idiom): a filter ringing out into silence otherwise wanders into denormal territory and multiplies its own CPU cost right when the music is quietest.
  • What is deliberately absent: fast-tanh approximations and a polyphase halfband resampler are both flagged in the file as candidates — and parked, because each changes output microscopically and the project's rule is that optimizations land only bit-identical or explicitly signed off.

Checkpoint

Trapezoidal integration turns the SVF's two integrators into a solvable linear system per sample — A-stability is where the modulation-proofness comes from, prewarping is where the tuning accuracy comes from, and the output mix is where morphing comes from, corner-exact by construction. Butterworth spread keeps cascades flat; resonance sharpens one section; the driven circuit's tanh placement makes bounded self-oscillation a fixed point rather than a feature. The rest is bookkeeping — and the bookkeeping was benchmarked.

The nonlinear loop: ladder.h

The user-facing chapter promised self-oscillation in tune (8009 Hz measured for an 8 kHz cutoff), THD that walks from 0.5 % to 33 %, comp recovering exactly the passband resonance eats, and a measured 13.5 dB from oversampling. This appendix derives all of it. The SVF appendix built the trapezoidal machinery; here it is wrapped in four tanh saturators and a feedback loop supposed to go unstable — and the engineering is solving a loop that no longer solves on paper.

The stage: one pole, trapezoidal, prewarped

Each stage is the analog one-pole lowpass y' = ω·(x − y), discretized with the trapezoidal rule exactly as in the SVF. Once per cutoff change, update_derived computes

g   = tan(π · fc / fs_os)      the prewarped integrator gain
m_g = g / (1 + g)              the solved per-stage gain, G below

and the per-stage step (tpt) is the standard zero-delay one-pole:

v = (x − s) · G        y = v + s        s ← y + v   (= 2y − s)

For a lone lowpass the tan prewarp is a nicety. Here it is the tuning system: the filter's oscillation frequency is set by where the stages put their phase, so pole mis-placement becomes pitch error — the failure of the classic Stilson/Smith ladder in the top octaves.

The loop: why the magic number is four

Four stages in series, global negative feedback: the stage-1 input is L − m_k·y4 (L the driven input, m_k = 4.0 * resonance). The 4 is the linear loop's oscillation threshold. At the cutoff a one-pole has response 1/(1 + j): magnitude 1/√2, phase −45°. Four in series:

|H⁴(fc)| = (1/√2)⁴ = 1/4        ∠H⁴(fc) = 4 · (−45°) = −180°

The input subtraction supplies the other 180°, so at fc — and only there — the loop phase is 360°. Barkhausen: oscillation begins at unity loop gain,

k · 1/4 = 1        ⇒        k = 4

So resonance = 1.0 (k = 4) is the mathematical edge, oscillation happens at the tuned cutoff, and self-oscillation frequency is the tuning test: the notebook measures 1000.2 Hz for a 1 kHz cutoff (0.02 % error) and 8009.0 Hz for 8 kHz (0.11 %) — the prewarp holding at the top of the keyboard, as promised.

The file allows k_res_max = 1.1, i.e. k = 4.4, "comfortably past self-oscillation." Past the edge the linear model diverges — but as amplitude grows, tanh's small-signal gain falls, the effective loop gain sags back toward 4, and the oscillation parks where they balance: the SVF driven circuit's fixed-point argument, no clipper needed. The notebook's oscillation (resonance 1.08) peaks near |y| = 0.10; the kernel test holds five seconds at k = 4.4 finite, under 2.0 peak, RMS steady within a 0.7–1.4× band. An all-zero state also solves the equations — hence the header's advice to ping it.

The ladder as a diagram: four tanh one-pole stages wrapped by the k = 4·resonance feedback, with comp and the pole-mix taps

The file as a schematic: the Barkhausen condition lives at the red tap.

The algebraic loop, solved linearly first

Zero-delay feedback through four stages means y4 depends on the stage-1 input, which depends on y4. Linearly, substitution closes it: chain the linear stage form y = G·x + B·s (B = 1 − G, s the held state) through all four with u = L − k·y4,

y1 = G·u + B·s1
y2 = G²·u + G·B·s1 + B·s2
y3 = G³·u + G²·B·s1 + G·B·s2 + B·s3
y4 = G⁴·u + G³·B·s1 + G²·B·s2 + G·B·s3 + B·s4

then name the state-only part S and solve:

S  = G³·B·s1 + G²·B·s2 + G·B·s3 + B·s4
y4 = G⁴·(L − k·y4) + S        ⇒        y4 = (G⁴·L + S) / (1 + k·G⁴)

That last expression is predict_linear verbatim — the code's (G2*G2*L + S) / (1.0 + m_k*G2*G2) with G2 = G*G and the same four-term S over m_s1..m_s4. For the linear ladder it is exact — the four-stage analog of the SVF's a1/a2 solve.

The saturators, and the one-pass commit

With tanh in every stage, the honest loop equation y4 = F(L − k·y4) has no closed form. The file ships two answers.

solver_fast (default) is Huovilainen-flavored prediction-correction: compute predict_linear(L) as if the saturators weren't there, then run the saturating stages once with that feedback value and commit (core):

t0 = sat(L − m_k·y4_est)
y1 = tpt(m_s1, t0, G),   y2 = tpt(m_s2, sat(y1), G),   ... y4 likewise

The committed y4 is not the y4_est the feedback used — that mismatch is the method's error. When the signal is small, tanh is the identity and the prediction is exact: the linear filter is recovered in the limit. The error grows only with how far tanh bends over one sample's state change — so drive × resonance is the failure axis, and oversampling (2× default) doubles as accuracy: it shrinks the per-step change being predicted.

solver_exact solves the true loop: Newton iteration on F(g) = y4_trial(L, g) − g, where y4_trial evaluates the four saturating stages for a guessed feedback value without touching state. Seeded by the linear prediction, clamped to ±3 (a tanh-bounded loop cannot park a fixed point far outside ±1), with a numerical derivative that falls back to the seed if it degenerates, at most 12 iterations to a 1e-12 residual. The commit reuses the same core path — with the converged g it reproduces the trial values while tpt advances the states. One code path, two accuracies.

How different are they? The kernel test renders both at drive 3 dB, resonance 0.5 and pins the maximum sample difference below 0.01; the stress test (drive 24 dB, resonance 1.1, asym 1.0) asks only that solver_exact stay finite and bounded. Audibly identical until drive and resonance are pushed — and solver_fast costs one saturated pass where Newton can cost dozens of trial evaluations per sub-sample.

asym: moving the operating point

Real ladder transistors don't match, so real stages don't saturate symmetrically. The model is an operating-point shift in every stage:

m_sat_bias = 0.3 · asym
sat(v) = tanh(v + m_sat_bias) − m_sat_dc      m_sat_dc = tanh(m_sat_bias)

The subtraction keeps sat(0) = 0 exactly — silence in, silence out. Expand tanh about the bias: a curvature term −tanh(b)·sech²(b)·v² appears only when b ≠ 0, and a v² term generates second harmonic and DC. The notebook measures the driven 2nd harmonic at −155.8 dB relative to the fundamental at asym 0 (numerical noise — tanh is odd), rising to −18.6 dB at asym 0.6. The DC is the rectifying side of the same v² term; the header owns it honestly and delegates to tap.dcblock~. Drive's own numbers: THD 0.54 / 3.45 / 16.50 / 33.07 % at 0 / 8 / 16 / 24 dB (notebook) — odd harmonics only, until asym says so.

comp: the passband bargain, quantified

At DC every stage passes unity; the closed linear loop gives

y4(DC) = L − k·y4(DC)        ⇒        y4(DC) = L / (1 + k)

Resonance eats the passband by exactly 1/(1+k). At resonance 0.9, k = 3.6: predicted 20·log10(1/4.6) = −13.3 dB; the notebook measures −13.2 dB. The compensation is a pre-gain (update_derived):

m_in_gain = 10^(drive/20) · (1 + comp·m_k)

At comp = 1 the input is multiplied by (1 + k) and DC gain returns to exactly unity — measured +0.0 dB — with a linear blend below. One honest note: the compensation multiplies the input before the saturators, so high comp at high resonance also leans harder on the tanh stages — like turning up the level into hardware; authentic, not a linear post-trim.

Pole mixing: the Xpander table

core returns a fixed weighted sum over the taps [t0, y1, y2, y3, y4] (k_c_mix). In the linear small-signal limit each tap is a power of the one-pole response H applied to u, so the mixes are polynomial algebra:

lp12:  y2                        =  H²·u
hp12:  t0 − 2·y1 + y2            =  (1 − H)²·u        weights {1,−2,1,0,0}
hp24:  (1 − H)⁴                  →  {1,−4,6,−4,1}      alternating binomial
bp12:  2·(y1 − y2) = 2·H(1−H)·u  →  {0,2,−2,0,0}
bp24:  4·H²(1−H)²                →  {0,0,4,−8,4}

The binomial rows are literally (1−H)ⁿ expanded. The bandpass factors are unity-gain normalizers: at fc, |H| = |1−H| = 1/√2 with phases ∓45°, so H(1−H) has magnitude 1/2 and phase 0 — the 2 (and 4 for its square) restore 0 dB at center. Measured high-side slopes: 23.4 dB/oct for lp24 (want 24), 11.7 for lp12 (want 12). Two caveats, both inherited from the analog original: under saturation the taps carry distortion products and the algebra is approximate (the header says so), and the feedback is always the full four-pole loop — a "12 dB" mode is a two-pole slope riding four-pole resonance, exactly as in an Xpander.

Oversampling: paying for tanh honestly

tanh generates harmonics without limit; above Nyquist they fold back inharmonically. run is the classic chain (the tap.verb~ pattern, self-contained per house rule): zero-stuff by the factor, scaling the retained sample by m_os to preserve passband gain; 4th-order Butterworth anti-imaging at 0.45 of the original Nyquist (fc_norm = 0.45/m_os); the nonlinear core at the high rate; a matching anti-alias Butterworth; decimation by keeping the last filtered sub-sample. Each Butterworth is two RBJ biquads at the textbook Q pair 0.54119610 / 1.30656296 = 1/(2·cos(π/8)), 1/(2·cos(3π/8)). Measured on a hard-driven 5 kHz tone: non-harmonic (alias) energy −30.1 dB at 1×, −43.7 dB at 4× — the promised 13.5 dB. The clamp fc ≤ 0.49·fs_os keeps 20 kHz legal at every factor.

The engineering ledger

  • One derived tier, not two. Unlike the SVF's split shape/cutoff caches, update_derived recomputes everything whenever anything moves (m_derived_dirty stays set while m_ramps_active > 0). Nearly every derived value reads several parameters (m_in_gain: drive, comp, and resonance) — a finer split buys little.
  • The signal-rate cutoff path re-dirties deliberately. process(x, cutoff_hz) recomputes for the override, then sets m_derived_dirty = true — "the cached G belongs to the override, not the parameter" — so the message-rate path never serves a stale one.
  • Ramps everywhere, counted. Every parameter rides a per-sample linear ramp (20 ms default); m_ramps_active makes idle one integer test. The kernel test bounds the worst sample jump through a 100 ms preset recall and a per-sample 500→6000 Hz sweep — click-free is asserted, not assumed.
  • Preset morph in the kernel. 16 slots; recall_preset is ramp_to on all six parameters with a shared duration — as safe as any motion.
  • Anti-denormal on the stage states (anti_denormal, the tap.comb~ 1e-15 idiom), inside tpt — a ringing-out filter otherwise decays into denormals and multiplies its own CPU cost.
  • Newton is guarded, not trusted. Seed clamp, derivative fallback, iteration cap: solve_exact cannot NaN or hang, only degrade toward solver_fast.
  • Allocation-free after prepare(); setters are plain stores into ramp targets, safe from the message thread while audio runs.

Checkpoint

Four trapezoidal one-poles put −180° and gain 1/4 at the prewarped cutoff; negative feedback makes k = 4 the oscillation threshold, which is why resonance is calibrated in quarters of k and self-oscillation lands on pitch (8009 Hz for 8 kHz, measured). The linear loop solves in closed form; the saturating loop is predicted linearly and committed through the tanh stages once — exact in the small-signal limit, backstopped by a guarded Newton solver. asym shifts the tanh operating point, comp pre-multiplies away the derived 1/(1+k) droop, the Xpander table is binomial algebra over the taps, and the oversampling chain pays tanh's alias bill with a measured 13.5 dB. Every number in the user-facing chapter traces to a line in this file.

The master phase and its corrections: vco.h

The user-facing chapter promised an oscillator whose folded harmonics sit ~47 dB down, whose analog section is "exactly zero by default" with a seed that works like a serial number, and whose FM survives through zero. Each claim is a theorem about this file or a measurement of it. This appendix derives the corrections — polyBLEP, the leaky triangle, the sync patch — then the analog-character section, including the one place where honest analysis contradicted intuition and the tests were written to match.

One phase, many readings

There is a single accumulator, m_phase ∈ [0,1), advanced once per sample in step:

f_eff = base_hz · 2^(cents/1200) + fm_hz      (pitch is exponential,
dt    = f_eff / m_sr,  clamped to ±0.49        FM is linear, in Hz)
adt   = max(|dt|, 1e-8)

cents collects detune, drift, jitter, the per-unit tolerance offset, and track — everything musical multiplies; only FM adds. Every waveform is a reading of the same phase: sine through sin, saw as 2p−1, pulse as a comparison against pw, triangle as an integral. The shape morph crossfades adjacent readings of one phase, so it can never produce a discontinuity the phase itself doesn't have. The problem is entirely the discontinuities.

The VCO as a diagram: frequency sum into the master phase accumulator, fanning to the four waveform readings and the shape crossfade

The fan-out the chapter title promises: every waveform is a reading of the same φ.

The residual: deriving poly_blep

A naive saw jumps by −2 at the wrap; a step's spectrum falls at only 6 dB/oct, so its harmonics march past Nyquist and fold back inharmonic. The ideal fix is a band-limited step — the integral of a sinc. The polyBLEP observation: the band-limited step differs from the naive step only near the edge, so instead of storing sinc-integral tables (minBLEP), approximate the difference with a polynomial. Take the crudest kernel one sample wide per side — a unit-area triangle, b(τ) = 1 − |τ| on τ ∈ [−1, 1] with τ in samples from the discontinuity — integrate, subtract the step:

s_bl(τ) = (τ+1)²/2             τ ∈ [−1, 0]
s_bl(τ) = 1 − (1−τ)²/2         τ ∈ [0, 1]

r(τ) = s_bl − s_naive:
r(τ) = (τ+1)²/2                τ ∈ [−1, 0)      (before the edge)
r(τ) = −(1−τ)²/2               τ ∈ [0, 1]       (after the edge)

Scale by 2 — the saw's wrap step — and these are exactly the file's poly_blep branches: just after the wrap (t < dt, normalized t/dt = τ), t + t − t*t − 1 = −(1−τ)² = 2·r; just before it (t > 1.0 − dt, normalized (t−1)/dt = τ ∈ [−1,0)), t*t + t + t + 1 = (τ+1)² = 2·r. Three properties fall out of the derivation:

  • Continuity at the window edges. r(−1) = r(1) = 0: the correction fades in and out without new discontinuities of its own.
  • The midpoint property. r(0⁻) = +½, r(0⁺) = −½: the corrected edge passes through the middle of the jump, as a true band-limited step does.
  • "±1 sample" precisely. A sample lands in a branch iff its phase is within dt of the wrap, and phase moves dt per sample — exactly the last sample before and the first after each edge are touched, ever; the scope trace still looks like a saw.

The triangle kernel approximates the sinc (its spectrum is sinc², not a brickwall), so suppression is finite and measurable: the notebook drives a 3951 Hz saw, whose 13th harmonic folds to 3364 Hz, and measures it at −26.7 dB naive versus −74.2 dB with polyBLEP — 47.5 dB of suppression. The notebook's sample-level zoom shows the mechanism: the corrected saw passes through +0.588 and −0.856 on its way down, where the naive saw jumps in a single step.

Saw, pulse, and the second BLEP

saw_at is the reading minus the residual: 2·bent(p,·) − 1 − poly_blep(p, dt) (subtracted: the wrap step is −2, poly_blep is normalized for +2). pulse_at is ±1 with two edges: the rising edge at the wrap (step +2, residual added) and the falling edge at p = pw (step −2, residual subtracted) — the latter evaluated at wrap01(p − pw), re-centering the phase coordinate so that edge sits at zero of its own window and the same branches apply. Calibration is pinned by measurement: a bipolar pulse at duty d must average 2d − 1, and the notebook measures −0.800 / −0.500 / +0.000 at 10 / 25 / 50 %; the kernel test holds the 25 % mean within ±0.03.

The triangle: integrate, but leak

A triangle is the integral of a square — the classic analog trick. A ±1 square at frequency f forces slope ±4f (2 units in half a period 1/(2f)), so the per-sample increment is ±4·f/fs = ±4·dt, which is tri_tick's scaling exactly:

m_tri_state = 0.999 · m_tri_state + 4.0 · adt · sq

giving peak ±1 with no post-normalization. The 0.999 is the honesty tax: the BLEP-corrected square's samples do not sum to exactly zero per period (the two edges land at different sub-sample positions, so their corrections don't cancel — and tri_pw skew under imperfect makes the imbalance deliberate), and a pure integrator would ramp that residue to infinity. The leak turns the integrator into a one-pole highpass with corner fs·(1 − 0.999)/2π ≈ 7.6 Hz at 48 kHz — far below any audible fundamental, high enough to hold DC bounded.

Why the integrated square is correctly antialiased: integration multiplies the spectrum by 1/ω, −6 dB/oct. The square's already-suppressed alias residual was generated near Nyquist, where 1/ω is smallest, so integrating a BLEP square improves the alias-to-harmonic ratio — the correction gets cheaper exactly where the waveform gets harder, which is why this hardware trick survives digitally intact.

Through-zero FM

Because FM adds in Hz after the exponential pitch math, dt can go negative and the phase genuinely runs backward — that is all "through zero" means, and why the sidebands stay coherent when the modulation swings past the carrier. Two guards make it safe. The BLEP windows use adt = |dt|: a window is a duration, one sample each side of an edge, whichever way the phase travels (with a 1e-8 floor so the t /= dt normalization survives a frozen phase). And dt is clamped to ±0.49: at |dt| ≥ 0.5 the window tests t < dt and t > 1 − dt would overlap and every sample would be "at an edge" — the clamp keeps the effective frequency below Nyquist, where the model means anything at all. Measured: a 500 Hz sine under ±900 Hz of FM at a 100 Hz rate — depth past the carrier, genuinely through zero — puts its sideband at −13.7 dB with −158.3 dB between the lines (144.5 dB of contrast), bounded at |y| = 1.00.

Hard sync, one-sided

A rising zero crossing on the sync input (m_sync_prev ≤ 0, sync > 0) resets the phase. Linear interpolation locates the crossing inside the sample:

frac    = m_sync_prev / (m_sync_prev − sync)      ∈ [0, 1)
m_phase = wrap01((1.0 − frac) · dt)

— the phase restarts from zero at the crossing and accumulates only the remaining fraction of the sample, so sync pitch is sub-sample accurate (the notebook's synced slave measures periodic at 110.1 Hz against a 110 Hz master). The reset is still a discontinuity of size d = waveform_out_peek(p_old, …) − waveform_out_peek(wrap01(p_new), …), sized on the morphed waveform without advancing the triangle integrator:

x = 1.0 − frac        correction += d · 0.5 · x²

This is a first-order polynomial BLEP, honestly cruder than the saw's: one-sided, because the pre-reset sample is already output when the edge arrives — a reset cannot be predicted — and a one-sided patch can never reproduce the full band-limited edge (the midpoint property needed both sides). d·½x² is the triangle-kernel residual for a step landing x into a sample; the code feeds it x = 1 − frac, the elapsed fraction since the crossing, so its weighting runs opposite to the two-point post branch (½·frac²) — at the first-order accuracy a one-sided patch can claim, both are O(d) click reducers vanishing at one end of the window. The header flags minBLEP tables as the wholesale upgrade; a m_pending slot, read and cleared each sample but never written, is scaffolding for the second correction sample it would need.

The analog section, derived (2026-07)

Two time scales of pitch noise. tick_drift is sample-and-hold noise redrawn every m_sr / 2 samples (~2 Hz) smoothed by a one-pole with a = 1 − exp(−2π·0.5/m_sr) — the exact discrete step of a 0.5 Hz lowpass. tick_jitter is the same structure at ~80 Hz through ~40 Hz: the fast companion, trembling where drift strolls. Both are depth-scaled in cents into the pitch path. Measured: relative period spread 2.01×10⁻⁷ at jitter 0 versus 2.74×10⁻³ at 10 cents — four orders of magnitude of micro-instability, still under the test's 0.02 ceiling.

The bent ramp, honestly. imperfect bows the saw via bent(p, bend) = p + bend·p·(1−p), bend = 0.35·imp·m_tol_curve. The parabola vanishes at both endpoints, so the wrap step stays exactly 2 and the BLEP stays correctly sized — why the bend lives inside the ramp reading. Now the honest part. In x = p − ½ the saw 2p−1 = 2x is odd (a sine series) while the parabola p(1−p) = ¼ − x² is even (a cosine series): the bend's Fourier content is in quadrature with the saw's own components. Harmonic k gains an orthogonal part of relative size bend/(πk) that moves its magnitude only at second order — about 0.05 dB at k = 1 for the maximal bend of 0.35, per √(1 + (bend/πk)²): a scope-obvious shark fin, almost no harmonic-magnitude shift. Discovered by measurement — and the kernel test matches the truth: it asserts the waveform bow (interior deviation > 0.03 at imperfect 1) and makes no harmonic-magnitude claim.

Where the spectral work is actually done. The reset corner rounds through a one-pole whose cutoff closes from ~22 kHz toward ~8 kHz (fc = min(22000 − 14000·imp, 0.45·m_sr), coefficient cached against m_round_imp): measured, the saw's 40th harmonic (17.6 kHz) is 6.4 dB quieter at imperfect 1; the test requires > 4 dB. The triangle skews via tri_pw = clamp(0.5 + imp·m_tol_tri·0.01, 0.05, 0.95) — duty asymmetry in the integrated square is even harmonics — measured: triangle h2 rises from −185.5 dB (numerically absent) to −34.1 dB at imperfect 0.8. The sine reads a mildly bent phase (bent(p, 0.5·bend)), the pulse width takes a static offset up to ±1.5 %, the whole unit a pitch offset up to ±2 cents (m_tol_cents).

Which unit you own. The tolerances come from a separate stream: compute_tolerances hashes the seed (m_seed * 2654435761u + 12345u) into its own local LCG, never touching the runtime m_rng. The contract: clear() resets m_rng = m_seed and all noise state but does not re-roll tolerances — resetting the oscillator must never change which unit off the production line you own; only set_seed re-rolls, because changing the serial number is changing the unit. Every tolerance is scaled by imperfect at use, yielding the contract the section rests on: at imperfect 0, every seed is bit-identical to the ideal oscillator. The kernel test renders seeds 7 and 8 and requires ya == yb — exact equality over 24000 samples — and the notebook confirms it; conversely, with drift 20 seeds 7 and 8 diverge by up to 0.183 (measured) while the same seed renders bit-identically, and the test pins that at imperfect 0.6 different seeds are audibly different units.

track.

cents += track · log₂(base_hz / 440.0)

A V/oct converter's calibration error grows linearly in octaves from its trim point; this is that line, exact at 440 Hz by construction (log₂ 1 = 0). Measured at track 5: −15.0 / −10.0 / −5.0 / +0.0 / +5.0 / +10.0 / +15.0 cents across −3…+3 octaves; the test holds the trim point under 1 cent, ±3 octaves within 2 cents of ±15.

The engineering ledger

  • Determinism is structural. All randomness flows from one 32-bit LCG (1664525 / 1013904223) seeded by m_seed (0 remapped to 1); no wall clock, no std::random — renders, tests, and mc. stacks reproduce bit-for-bit, and the jitter test pins same-seed bit-identity.
  • Off means exactly off. tick_drift/tick_jitter return before consuming the RNG at zero depth — a default-configured oscillator never advances m_rng, so different seeds render identically until a stochastic feature is engaged. The corner-rounding pole is gated on imp > 0.0, its state primed while bypassed (m_round_lp = y) so engaging imperfect mid-note is click-free.
  • The triangle integrator ticks only when the morph needs it. waveform_out short-circuits the crossfade endpoints (a <= 0.0), and tri_tick is stateful — skipping it when unused is a cost saving and a correctness rule (parking at pure saw must not silently integrate); sync sizing uses waveform_out_peek for the same reason.
  • Clamps with reasons: dt at ±0.49 (window overlap / Nyquist), adt floored at 1e-8 (division in poly_blep), tri_pw in [0.05, 0.95] and pw in [0.01, 0.99] (an edge pair must stay two distinct windows).
  • The house frame: per-sample linear ramps with an active-count fast path, 16 preset slots morphable over time, allocation-free processing, setters safe while audio runs — the same bones as ladder.h and svf.h, so the wrapper stays a shim.

Checkpoint

One master phase; every waveform is a reading of it, and every reading's discontinuity gets the residual of a triangle-kernel band-limited step — two samples per edge, continuous at its window boundaries, a measured 47.5 dB of alias suppression at the folded 13th harmonic. The triangle integrates the corrected square (slope ±4f, hence 4·adt) with a 0.999 leak; FM adds in Hz so the phase can run backward, adt keeping the windows directionless; sync resets with sub-sample accuracy and patches the step one-sidedly because resets can't be predicted — minBLEP is the flagged upgrade. The analog section is derived noise at two time scales, a quadrature-honest bent ramp, a rounding pole, a skewed duty, and a calibration line exact at A440 — all drawn from a tolerance stream clear() never touches, all scaled by imperfect, and all provably absent at zero: the ideal oscillator is a test-pinned invariant, not a default setting.

Detector, law, and a borrowed filter: autowah.h

The user-facing chapter (The pedal that listens) made three measurable claims: the sweep law matches its design to 0.000 cents, the follower's timing is an honest RC discharge, and sensitivity at the floor turns the object into a truly fixed filter. This appendix derives each, starting from the decision that tap.autowah~ is mostly not a new filter.

The composition decision: don't write a second SVF

The Snow White's core is an LM13700 OTA state-variable filter swept by an envelope. TapTools already ships an SVF kernel whose defining property — proved in the SVF appendix — is A-stability under per-sample cutoff modulation. An envelope-swept filter moves its cutoff every sample by construction; the property the wah needs most is exactly the one svf.h already guarantees by theorem. So wah_filter owns a tap::tools::svf::svf_filter member (m_svf) and drives it through the signal-rate path, m_svf.tick(m_cutoff) then m_svf.process(0, x), once per sample. The house rule makes this legal: objects under source/projects/ stay self-contained, but inside the kernel repo sharing between kernels is encouraged — autowah.h simply #include "svf.h".

Composition buys something subtler than saved code: the corner-identity argument. When the envelope is off, the wah is a bare SVF, and that is a testable equation rather than a resemblance. The kernel test ("sensitivity at the floor is the cocked-wah: bit-close to a bare svf at bias") runs the wah at sensitivity −60 dB, bias 800 Hz, resonance 0.7 against a separately constructed svf::svf_filter fed ref.process(x, 800.0), and requires maxerr < 1e-12 over half a second of signal. The wet paths are arithmetic-identical — same tick(cutoff) entry, same clamps, same solve — so why 10⁻¹² and not ==? The dry leak: at mix = 100 the mix angle is θ = π/2, and while sin(π/2) is exactly 1.0 in doubles, cos(π/2) rounds to 6.123×10⁻¹⁷, so the output carries a 6×10⁻¹⁷·x dry residue. The tolerance covers one ulp-scale cosine, nothing else. Resonance meaning is shared the same way — the wah's 0..1 knob goes through the SVF's own q_from_resonance mapping, so "resonance 0.7" means the same Q in both objects.

The autowah composition as a diagram: detector chain and sweep law around the boxed borrowed svf_filter member

The composition decision, drawn: everything amber is this file; the blue box is svf.h, borrowed intact.

The detector: gain, rectifier, follower

The detector chain in process() is three lines:

driven = key · m_sens_gain                           (dB → linear input gain)
rect   = |driven|            (full-wave, default)
       or max(driven, 0)     (half-wave, the traced single-diode topology)
m_env += coef · (rect − m_env)                       (one-pole follower)

The sensitivity floor is a contract, not a clamp. update_derived maps the dB knob as

m_sens_gain = 0                    if sens_db ≤ −60 dB
            = 10^(sens_db / 20)    otherwise

−60 dB is not "very quiet" (that would be gain 0.001); it is exactly zero. With m_sens_gain = 0 the rectifier output is identically 0, m_env decays to 0 and stays there, m_sweep = tanh(0) = 0, and map_cutoff(0) returns m_bias exactly — the test asserts w.cutoff_hz() == 250.0 with ==. That is what makes the pedal's secondary "cocked wah" mode a true fixed filter rather than an approximately-fixed one that still breathes a few cents with the input. Factory slot 3 is that voicing as data.

The follower coefficient is the exact RC discretization. The analog detector is a capacitor charged toward the rectified signal: env′ = (rect − env)/τ. Solving that ODE exactly over one sample period T = 1/fs gives env[n] = rect + (env[n−1] − rect)·e^(−T/τ), which rearranges to the code's recurrence with

a = 1 − e^(−1/(τ·fs))        (the file: m_attack_coef = 1 − exp(−1000/(ms·m_sr)))

After N = τ·fs samples of a step, the remaining error is (1−a)^N = e^(−N/(τ·fs)) = e^(−1): the envelope reaches 63.2% in exactly τ. This is not the cheap approximation a ≈ 1/(τ·fs); the exponential form makes the ms parameters honest at any rate. The notebook measured it: attack set to 2.0 ms reaches 63% in 1.94 ms; decay set to 250 ms falls to 36.8% in 256 ms; and a log-domain fit of the release is a pure exponential with τ = 252 ms and residual σ = 0.004 — an RC discharge, like the hardware. The attack/release asymmetry is one branch, coef = (rect > m_env) ? m_attack_coef : m_decay_coef — the diode charges the cap through one resistance and lets it bleed through another.

Full-wave default, half-wave option. The traced hardware detector is a single diode: it charges only on positive half-cycles, so the envelope droops between charges at the signal's fundamental. Full-wave rectification charges twice per cycle and halves the gaps. The follower cannot filter this out without also slowing the response, so the ripple rides the envelope and frequency-modulates the cutoff — the hardware's "sweep-rate ripple." The notebook quantified the A/B on a 110 Hz tone (decay 60 ms): settled envelope ripple (std/mean) 0.7% full-wave vs 2.6% half-wave. The kernel defaults to the cleaner full-wave and keeps half-wave selectable (set_rectifier()) because the flavor question is a hardware-listening question, not a math question — it waits for the calibration pass.

The sweep law, and where it is honest about ignorance

m_sweep  = tanh(k_env_knee · m_env)                     k_env_knee = 1.5
m_cutoff = m_bias · 2^(m_sweep · m_range)               clamped to [20 Hz, min(20 kHz, 0.45·fs)]

Three deliberate choices:

(a) Exponential in Hz. Equal envelope increments move the cutoff by equal octaves, which is how a sweep sounds uniform — pitch perception is log-frequency. The honest caveat lives in the header and in map_cutoff()'s own comment: the LM13700's frequency is linear in control current, so the pedal's true law hinges on the BJT stage that converts the envelope voltage into that current — plausibly exponential (a BJT's collector current is exponential in V_BE), but not yet measured. That is why the law is one isolated function: if calibration finds a linear V→I driver, map_cutoff becomes bias + sweep · span_hz and nothing else in the kernel changes.

(b) The tanh soft knee. Without it, hard playing would pin sweep at a clamp rail — a hard corner in the control trajectory, audible as the filter slamming its ceiling. With it, the ceiling is approached asymptotically. The arithmetic at the defaults (bias 250 Hz, range 3.3 octaves, sensitivity 0 dB):

full-scale DC key → m_env → 1
m_sweep → tanh(1.5) = 0.905
m_cutoff → 250 · 2^(0.905 · 3.3) ≈ 1982 Hz

— about 2 kHz, under the asymptotic rail 250·2^3.3 ≈ 2462 Hz, which itself matches the hardware's published 250–2500 Hz span. The unit test pins the settled cutoff into (1800, 2100) Hz and separately drives an absurd +24 dB sensitivity into an 8× full-scale key to confirm the cutoff saturates at the ceiling (reaching > 99% of it) instead of running away. range is signed — negative sweeps down from bias, a deliberate extension the pedal never had.

(c) Measured. The validation notebook swept the envelope range and compared measured cutoff against the designed curve: max error 0.000 cents. The law in the code is the law on paper; when hardware recordings arrive, any disagreement is a fact about the model choice, not about the implementation.

One filter, two owners: the forwarding discipline

Both kernels ship per-sample parameter ramps. Run both and every set would be smoothed twice — lagged, and worse, shaped (a ramp of a ramp is not a ramp). So prepare() declares a single owner:

m_svf.set_smooth_ms(0.0);   // this kernel owns all smoothing; svf setters snap

The wah's own ramp array smooths every audible parameter; the composed SVF's setters snap instantly to whatever the wah forwards. And forwarding is change-gated: update_derived pushes set_resonance / set_drive / set_circuit only when the cached values (m_svf_resonance, m_svf_drive, m_svf_circuit, seeded to −1 to force the first forward) actually differ. The reason is the SVF's two-tier update from its own appendix: those setters dirty the shape tier (damping, mix weights, drive gain — a pow and the mix logic). While any wah ramp is active, update_derived runs every sample; forwarding unconditionally would re-run the SVF's shape update every sample of every bias morph even though resonance never moved. Cutoff needs no gate at all — tick(cutoff_hz) lands in the SVF's cutoff cache, which recomputes the tan and solve constants only when the value differs.

The circuit switch

drive at 0 dB runs the SVF's clean linear circuit; anything above engages circuit_driven — tanh band-node limiting, 2× oversampled — as the optional OTA-flavored color stage. The switch is a threshold in update_derived:

circuit = (m_svf_drive > 1e-6) ? circuit_driven : circuit_clean

Why is switching circuits mid-stream acceptable? At the switch point drive ≈ 0 dB, so the driven circuit's input gain is 1 and tanh is near-identity at typical band-node levels — the two circuits compute nearly the same output, and the transition is benign. Honestly stated: near-identity, not identity. At high resonance the band node runs hot, tanh visibly bends, and the driven circuit also brings its oversampling path with it — so engaging drive from zero on a screaming resonant setting is a small audible step. The abuse test accepts this trade explicitly: resonance 1.0 plus max drive on square-wave bursts must stay bounded and finite (the SVF's bounded self-oscillation doing its job), not polite.

Output staging is the equal-power crossfade shared with tap.crossfade~: θ = mix·π/200, m_dry_gain = cos θ · g, m_wet_gain = sin θ · g — the master gain g rides both paths, so gain never changes the balance.

The engineering ledger

  • Per-sample cost accounting. Settled steady state pays: one rectify, one follower multiply-add, one tanh (knee), one exp2 (law), the SVF solve, two mix multiplies. The pow(10, ·)/exp calls live only in update_derived, which runs per sample while ramps move and exactly once after they settle (m_derived_dirty re-arms only when m_ramps_active > 0). The real recurring cost is the SVF's cutoff-tier tan, paid whenever the envelope actually moved the cutoff — and skipped by the SVF's cutoff cache whenever it didn't (silence, or the cocked wah).
  • envelope() and cutoff_hz() exist for measurement. They read m_sweep and m_cutoff after the fact; the C ABI's taptools_wah_process(..., env_out, cutoff_out, n) taps them per sample, and the validation notebook's trace=True path — the ground truth its STFT peak-trajectory extractor was proven against (0.979 log-frequency correlation) — is built on exactly these two accessors.
  • The preset-morph engine is the GRM pattern (16 slots, the grm_comb/grm_pitchaccum house count): store_preset captures ramp targets (knob positions, never mid-ramp instantaneous values), recall_preset(slot, seconds) re-targets every ramp so a morph is just nine simultaneous ramps — re-targeting mid-morph stays continuous for free. The test walks a 100 ms morph and requires bias to move monotonically with no step larger than one ramp increment.
  • Factory slots are data, not code: four params structs (guitar / bass / slow swell / cocked wah) in slots 0–3. Changing a voicing after the hardware session edits numbers, not logic; the test pins them.
  • Structural switches (mode, rectifier) are not ramped or morphed — interpolating between rectifier topologies has no physical meaning.
  • Anti-denormal on the envelope (< 1e-15 → 0, the tap.comb~ idiom): a decaying exponential otherwise glides into denormals and multiplies its CPU cost during silence.
  • Sidechain by signature: process(x) is process(x, x); the wrapper's key inlet is the two-argument form. Single-channel by design — per-channel envelopes are the correct behavior under mc. wrapping.

The calibration pass, by construction

Every open hardware question maps to one isolated switch point: the sweep law is map_cutoff() (one function), the stock filter tap is m_mode's default (mode_lowpass, flagged in the header as inference), the detector topology is set_rectifier(), the knee is k_env_knee (one constant). The validation notebook is the instrument that will close them: its extractor recovers the swept peak from wet audio alone, is already calibrated against the kernel's own trajectories, and its last cell waits for snowwhite_*.wav. When the pedal arrives, disagreements land on named constants — not on a rewrite.

Checkpoint

The wah is a detector and a law in front of a borrowed filter. Composing svf_filter puts the per-sample-modulation stability where it is already proven, and makes "sensitivity off equals a bare SVF" an identity checked to 10⁻¹² (the gap being one rounded cosine). The follower coefficient 1 − e^(−1/(τ·fs)) is the exact RC step — 63.2% in τ by algebra, 1.94 ms measured for 2.0 set. The sweep law is exponential-in-Hz through a tanh knee that turns hard playing into asymptotic approach (~2 kHz at the defaults, under the hardware's 2.5 kHz rail) — measured at 0.000 cents against design, and honestly provisional, isolated in one function until the real pedal votes.

Convolution without compromise: conv_engine.h

The user-facing chapter (Borrowed rooms) made a flat claim: tap.convolve~ is exact — not "high quality," exact — and its only cost is a latency of precisely blocksize samples. Claims like that are either algebra or advertising. This appendix does the algebra: why the convolution is partitioned at all, why overlap-save, why the latency is exactly one partition, and how an impulse response can be replaced mid-performance without the audio thread ever seeing a torn table.

Why partitioned: the cost triangle

Direct convolution of a stereo pair against an L-second IR at rate fs costs

cost_direct = L·fs MACs per output sample per path

— 48 000 multiply-adds per sample for one second of room at 48 kHz, times four paths for true stereo. Untenable. The classical fix is to convolve in the frequency domain: transform the whole IR once, multiply spectra, invert. But a single-FFT scheme cannot emit anything until it holds a full frame of input, so its latency equals the IR length — seconds of delay for a reverb. Also untenable.

Uniform partitioning takes the middle of the triangle. Split the IR into P equal partitions of B samples, h_j[k] = h[j·B + k]; then by linearity

h = Σ_j h_j delayed by j·B

Each partition is short enough that its convolution can be computed with a small FFT once per B-sample block, and the delays j·B are whole blocks — which, we will see, cost nothing but indexing. FFT economics, latency of one partition. This is the standard engine of the genre for a reason.

Overlap-save: the framing, derived

The engine convolves each partition by circular convolution over an FFT of size m_fftsize = 2·m_block — size 2B for partition size B. Circular and linear convolution are not the same thing; the design question is which output samples of the circular product are also the linear ones.

Take the frame the code actually builds in process_block():

frame_m = [ block_{m−1} ; block_m ]        (m_fre[j] = m_prev[ch][j]; m_fre[B+j] = m_inblk[ch][j])

and a partition h zero-padded from B to 2B. The circular convolution is

y_circ[n] = Σ_{k=0}^{B−1} h[k] · frame[(n − k) mod 2B]

For n in the second half, n ∈ [B, 2B), and k < B, the index n − k stays in [1, 2B): the mod never wraps, so y_circ[n] equals the linear convolution of h with the input stream at that time. For n < B the mod does wrap, splicing in samples from the frame's far end — time aliasing. So each 2B-point product yields exactly B valid samples, the second half, and the code keeps precisely those:

m_outblk[oc][j] = m_are[m_block + j];   // overlap-save: discard the aliased first half

That is the whole scheme: hop by B, keep the clean half, discard the dirty half. The alternative, overlap-add, zero-pads each input block instead and sums overlapping output tails — equally exact in theory, but it carries a partial-sum accumulation buffer across block boundaries, one more piece of state to get right. Overlap-save's output block is finished the moment the IFFT returns: no summation state, no output windowing, no crossfading — there is nothing between the IFFT and the output buffer that could be inexact. Every step in the chain — framing (a copy), FFT and IFFT (the shared radix-2 in fft.h, whose inverse divides by N so a round trip reconstructs its input), and the multiply-accumulate — is exact linear algebra in double precision. The engine has no tuning parameters that trade accuracy for speed; its error budget is rounding noise, and the measurements below confirm that is all there is.

The frequency-domain delay line

The FDL as a diagram: the ring of past input spectra multiplied per bin against the partition spectra, accumulated, and inverse-transformed with the aliased half discarded

Why FFT cost is constant in IR length: the transforms bracket the structure, and only the MAC sees the partitions.

The delays j·B remain. Delaying partition j's contribution by j blocks is the same as convolving it with the input from j blocks ago — and the frame for block m − j has already been transformed. So the engine keeps a ring of past input spectra (the FDL, m_fdl_re/m_fdl_im, one per input channel) and forms the output spectrum as

Y_m[k] = Σ_{j=0}^{P−1} H_j[k] · X_{m−j}[k]

which is exactly the inner loop: slot = cur − p (mod m_max_parts), then a complex multiply-accumulate over all m_fftsize bins into m_are/m_aim. The consequence for cost is the design's payoff: one forward FFT per input channel and one inverse FFT per output channel per block, regardless of P. Growing the IR grows only the MAC. Per output sample, the MAC costs P·2B complex MACs / B samples = 2P complex MACs ≈ 8P real multiplies per path, versus P·B real MACs for direct convolution — a factor of B/8 (64× at B = 512), with the FFTs an O(log B) constant on top.

Latency is exactly B, by the framing

Follow one sample through process(): it is written into m_inblk at position m_pos, and the output handed back at that same call is read from m_outblk[m_pos] — a block computed when the previous block completed. When block m finishes, process_block() runs with a frame ending at the newest sample, and its B valid outputs are the linear convolution up to that sample; they are then dealt out during block m + 1. So output sample t carries y(t − B): the first partition's contribution to a block is computed from the block just gathered, not from anything older, and the delay is one partition — no more (the frame includes the newest sample) and no less (nothing can be emitted before a block is complete). The unit test pins both edges: the first B output samples are silence (the pre-roll), and an IR of δ at index 5 yields the input delayed by exactly B + 5. The notebook verified the latency at B = 64, 256, and 1024 — always exactly B.

Exact, measured

The notebook (executing the real engine through the C ABI) puts numbers on "exact": against a direct time-domain convolution of the same IR — 24 000 samples, B = 512, 47 partitions — the maximum difference is 3.13×10⁻¹². Across block sizes the outputs agree with the direct reference to 2.27×10⁻¹³ / 1.17×10⁻¹² / 1.76×10⁻¹² (B = 64/256/1024), and with each other, latency-removed, to 1.81×10⁻¹² — the block size is a CPU/latency dial with no audible existence. An impulse through the engine reconstructs the loaded IR to 5.5×10⁻¹⁴, and a synthetic 0.60 s-RT60 reverb measures back at 0.599 s. The unit test does the same job in CI with independent per-path IRs, deliberately awkward 10-sample process chunks that straddle block boundaries, and 10⁻⁹ tolerances.

True stereo: four paths, two FFTs

A stereo room is a 2×2 linear system, and the engine runs all of it:

out_l = in_l ∗ h_LL + in_r ∗ h_RL ;   out_r = in_l ∗ h_LR + in_r ∗ h_RR

with path = in_channel·2 + out_channel (0 = LL, 1 = LR, 2 = RL, 3 = RR). The economics are better than 4× mono: the two input FFTs are shared across all four paths, and each output channel needs one inverse — so a block costs 2 forward FFTs, 2 inverse FFTs, and 4 MAC passes. The notebook pins the routing: an impulse into L only emerges on R at exactly the cross-feed path's gain (0.600 expected, 0.600 measured) and, with the off-diagonal paths silent, R stays at 0.0 — cross-terms cannot hide in each other. The mapping from buffer~ channel count to paths (4+ = true stereo, 2 = dual mono, 1 = same room both sides) is wrapper policy; the engine only ever knows four pointers, any of which may be null for a silent path.

The atomic IR swap

Loading a room while the music plays is the one place this engine touches concurrency, and it is confined to a single atomic. The IR tables are double-buffered per path (m_ir_re[path][slot], slots 0/1). load_ir() — which runs off the audio thread; it is the expensive part, P analysis FFTs per path — writes only the inactive slot, then publishes:

m_slot_parts[inactive] = P;                          // written before the publish...
m_active.store(inactive, std::memory_order_release); // ...so (slot, P) stay consistent.

The perform loop does one acquire load of m_active per block. The release/acquire pair means that if the audio thread observes the new slot index, it also observes that slot's fully written spectra and its partition count — slot and P travel through one atomic, so there is no window where the loop MACs over half-written tables or over the wrong number of partitions. Until the store, the loop reads the old slot, which the loader never touches. The discipline is single-writer double-buffering: publishes are serialized through the wrapper's message path, and the just vacated slot is only rewritten by the next load.

Why does a swap settle in exactly one block? Because the FDL stores input spectra, not output. The first process_block() after the publish already renders the entire tail — all P partitions — from the new IR against the existing input history; the only samples that differ from a new-IR-from-the-start engine are the ones sitting in m_outblk, computed just before the swap. The unit test pins the settling (output equals the new IR's pure delay from shortly after the swap); the notebook pins it exactly: max |swapped − reference| after +1 block = 0.00×10⁰ — bit-identical to an engine that had the new IR from the start — with RMS continuity across the swap instant, 21.922 before, 22.147 just after, no dropout. Honest limit: the swap is a hard splice between two exact convolutions, click-free but discrete; the engine does not interpolate between rooms.

A related freebie of the input-side FDL: analysis of incoming audio happens before the has-IR check, so the delay line is warm even while no IR is loaded — load the first room mid-stream and its tail renders immediately from audio already played.

The engineering ledger

  • Uniform, not Gardner non-uniform, partitioning. Non-uniform schemes (short partitions first, growing later) can push latency below B for the same CPU, but they need multiple FFT sizes and a scheduler that spreads long-partition work across blocks — real complexity with real failure modes. The object's latency budget is satisfied by making B small (64 samples = 1.3 ms, verified exact above); complexity was not bought that nothing needed.
  • The shared FFT. fft.h is the in-house radix-2 Cooley–Tukey used by the whole spectral set; per its header it lived byte-identical inside conv_engine, tap.nr~, and tap.spectra~ before being consolidated at the kernel split. In-place, forward unscaled, inverse divides by N — round-trip exact by construction. No external FFT dependency, per the porting philosophy.
  • IR stored as float32, deliberately. A buffer~ holds 32-bit samples; the engine quantizes at load_ir (static_cast<double>(src[idx]) * scale) and computes in doubles thereafter. The notebook's direct-convolution reference casts its IR through float32 the same way — so the 10⁻¹² figures isolate the algorithm, not the source quantization the wrapper inherits from Max regardless.
  • Geometry only in configure(). Partition size and capacity determine every buffer, so reallocation happens only where the audio thread is idle — the wrapper calls it from dspsetup. clear() flushes running state with std::fill only, no reallocation, and is safe from a message handler; process() allocates nothing, ever (scratch spectra m_fre/m_fim/m_are/m_aim are preallocated and reused).
  • Capacity vs. length. The FDL ring is sized m_max_parts; a loaded IR uses P ≤ m_max_parts partitions (load_ir clamps), and the MAC runs over P only — a short room in a big engine costs a short room.
  • The deferred optimization, on the record. The spectra are stored and MAC'd full-complex; the input is real, so a half-spectrum (N/2 + 1 bins, Hermitian symmetry) form would halve both the MAC and the IR/FDL memory. The header flags it and parks it, under the same house rule the SVF appendix recorded: optimizations land bit-identical or explicitly signed off — and re-deriving the packing arithmetic is exactly the kind of change that gets signed off with a measurement, not slipped in.

Checkpoint

Partitioning splits the IR by linearity; overlap-save framing makes each 2B-point circular product yield B exactly-linear samples with nothing to window or crossfade; the frequency-domain delay line turns partition delays into ring indexing, so FFT cost is constant in IR length and only the MAC grows. Latency is one partition by the framing — measured at exactly B for every B tried — and exactness is measured at 10⁻¹²-and-below everywhere it can be probed. The one concurrent act, swapping rooms, rides a single release/acquire atomic over double-buffered tables, and settles bit-identically in one block because the delay line remembers input, not output. The compromises the genre usually accepts — approximate tails, block-size coloration, swap dropouts — are absent, and the measurements say so.

Ring time as the truth: grm_comb.h

The user-facing chapter made three claims that sound like marketing until you do the math: that a voice keeps its decay as its pitch sweeps, that warp stretches the partials while the fundamental stays in tune, and that phase at 100 cancels the even harmonics — exactly. This appendix derives all three the way the file was designed, then walks the code-level decisions. The behavioral claims below are pinned by the kernel scenarios in tap.5comb_tilde_test.cpp, which drive tap::tools::fivecomb::comb_bank directly (no Max in the loop); the few numbers outside the test suite are marked as measured on the kernel for this chapter.

A comb is a string

One voice is a delay of d samples fed back on itself. Ignoring the in-loop filters for a moment, the recursion the code implements (y = in + fb * ap_out, written back into the delay line) is:

y[n] = x[n] + fb · y[n − d]        h[n] = δ[n] + fb·δ[n−d] + fb²·δ[n−2d] + …

The impulse response is echoes every d samples, each scaled by another factor of fb — and echoes every 1/f seconds is a tone at f and its harmonics: a plucked string tuned to f = fs/d. The first kernel scenario pins the geometry: a 500 Hz voice at 48 kHz (d ≈ 96) puts its first three echoes at one period spacing, within a couple of samples of 96/192/288.

Ring time as the truth

The legacy object exposed fb directly. This file refuses to, and the reason is in the impulse response above. After t seconds the signal has made t·fs/d round trips, so the decay envelope is level(t) = fb^(t·fs/d).

Reverberation's standard yardstick is RT60, the time to fall 60 dB — a factor of 10^(−60/20) = 10⁻³. Set level(rt60) = 10⁻³ and solve:

fb^(rt60 · fs / d) = 10⁻³   ⇒   fb = 10^(−3·d / (rt60·fs))

which is character for character the file's line in update_derived(): m_fb[v] = min(pow(10.0, −3.0·d_total/(rt60·m_sr)), k_fb_max).

The res knob maps to rt60 on a log curve before this — rt60 = k_rt60_min · (k_rt60_max/k_rt60_min)^(res_eff/100), 20 ms at res → 0⁺ up to 100 s at res 100 — because equal knob travel should mean equal ratios of decay time, which a linear map to fb spectacularly is not (all the action of a raw-feedback comb lives in the last few percent of the knob).

Now the design consequence, which is the chapter title. Because fb is re-derived from the current delay every time update_derived() runs, the ring time is the invariant: sweep a voice's frequency and fb is silently re-solved to hold rt60 constant. A raw-feedback comb has it backwards — hold fb fixed and rt60 = −3d/(fs·log₁₀ fb) is proportional to d, so low notes ring 1/f longer and high notes choke. The "resonance maps to ring time" scenario pins the calibration: inverting the log curve for rt60 = 1 s gives res ≈ 45.93, and the measured tail drops close to the ideal 30 dB over a half-second window (the test accepts 22–40 dB; it lands near 32, the excess being upper partials that the interpolator and loop lowpass damp slightly faster).

Why the delay must be fractional

At 48 kHz a 440 Hz comb needs d = 48000/440 = 109.09 samples. Round to 109 and the voice plays 48000/109 = 440.37 Hz — about +1.4 cents. Worse than the absolute error: each of the five voices quantizes differently, so the carefully-tuned beating between voices (the point of a bank) is replaced by whatever the rounding produced. The legacy abstraction had integer delays and control-rate stepping; the file header names that as the main reason it never sounded like the GRM original.

So the tap is fractional — but not linear. Reading between samples with linear interpolation is a two-tap filter H(z) = (1−η) + η·z⁻¹ where η is the fractional part, with magnitude:

|H(ω)|² = 1 − 2η(1−η)(1 − cos ω)

— a lowpass whose damping depends on η (worst at η = ½, where |H| = cos(ω/2), a null at Nyquist). Two failure modes follow. Statically, this filter sits inside the loop: its droop is applied once per round trip and compounds into the ring, so two voices with different fractional parts get different brightness decay for free. Dynamically, a sweep cycles η through 0 → 1 repeatedly, so the loop's damping ripples at the sweep rate — audible dulling and level flutter. read_hermite() is the 4-point, 3rd-order Hermite (Catmull-Rom) interpolator instead: C¹-continuous, passband flat to far higher frequency, far weaker dependence on η. Its cost is the geometry constraint noted in the code — the youngest of its four points is one ahead of the base, so d must exceed 2 strictly, hence k_min_delay_samples = 2.5 and the ceiling f_ceil = min(k_freq_ceil_hz, m_sr / k_min_delay_samples).

The feedback chain, in order

Per sample, the loop path in comb_voice::process() is:

delayed = read_hermite(d_read)  →  one-pole lowpass  →  DC blocker
        →  warp allpass  →  × fb  →  + in  →  write

The comb voice ring as a diagram: delay line, Hermite read, then the feedback chain through lowpass, DC blocker, warp allpass, and the ring-time feedback gain back to the input sum, with the d/2 pickup tap branched to the output

The lowpass is the string's brightness decay: every round trip gets a little darker, highs first, like a real string. Its coefficient is exact — m_lp_a[v] = 1 − e^(−2π·fc/m_sr) — placing the −3 dB corner at fc by construction (the one-step discretization of an RC section). The in-file comment flags the deviation: tap.comb~ used hz·2/sr, which is not even the small-argument limit of the exact map (that would be 2π·fc/fs) — its actual corner lands near fc/π, a factor-of-three tuning error on a labeled frequency knob. Faithful porting stops where the parameter lies about its units.

The DC blocker (y = norm·(x − x1) + R·y1, R = k_dc_block_r = 0.999, norm = k_dc_block_norm = (1+R)/2, ~7 Hz corner at 48 kHz) replaces the legacy tap.comb~ hard ±1 autoclip — the file's most consequential retirement. The clipper existed to stop runaway; but a clipper in a resonant loop is a distortion stage, and at high resonance — precisely where the GRM sound lives — the legacy object audibly distorted. The modern argument: cap fb below unity (k_fb_max = 0.99999), kill the loop's DC transmission (the blocker's zero at z = 1), and the linear loop contracts — no limiter needed, so res 100 rings clean.

The norm factor is this chapter's own contribution, and the story is worth a paragraph. The raw blocker (1 − z⁻¹)/(1 − R·z⁻¹) is not passive: its magnitude peaks at 2/(1+R) ≈ 1.0005 toward Nyquist. While proving the contraction claim for the first edition of this chapter, the measurement came back false in one corner: with the loop lowpass wide open the product |H_lp·H_dc| crossed unity near 450 Hz, and a voice tuned there at res 100 — where fb saturates at the cap — measurably swelled at ~+0.2 dB per second. The fix is the normalization: scaling the blocker by (1+R)/2 pins its peak gain at exactly 1 (the zero at DC is untouched), so with the allpass at unit magnitude and the one-pole lowpass ≤ 1, the loop gain is bounded by fb alone and fb < 1 now really is the airtight inequality. A kernel test pins the formerly-failing corner: 450 Hz, res 100, lp at 20 kHz, twelve seconds of ring-out, decaying window over window.

The normalization has a side effect the file also pays for: the blocker now slightly attenuates each voice's fundamental (a few parts in 10⁴ at mid frequencies, more for very low voices), which — uncompensated — would shave the top off long ring times. So update_derived divides the RT60-derived fb by dc_block_gain(ω₀), the blocker's magnitude at the voice's fundamental — the same pay-the-fundamental-back philosophy as the warp compensation below, and clamped to k_fb_max so the contraction bound survives. The second new kernel scenario pins the payoff: an impulse-excited voice at res 50 measures its RT60 within 10 % of the map's 1.41 s target.

warp: dispersion, and paying the fundamental back

The allpass is the modern GRM Comb's character control. warp sets

m_ap_c = −k_warp_coef_max · warp/100        (c ∈ [−0.85, 0])

and inserts H(z) = (c + z⁻¹)/(1 + c·z⁻¹) into the loop — unit magnitude everywhere (it cannot alter the decay), pure phase. At c = 0 it degenerates to z⁻¹, an honest one-sample delay: warp 0 is exactly the harmonic Classic comb. Its phase delay in samples is what allpass_phase_delay() computes:

τ(ω) = [ atan2(sin ω, c + cos ω) − atan2(c·sin ω, 1 + c·cos ω) ] / ω

τ(0) = (1 − c)/(1 + c)      (the DC limit the w < 1e−9 branch returns)

For negative c, τ falls monotonically with frequency — at c = −0.85, from (1.85/0.15) ≈ 12.3 samples at DC down to 1 sample at Nyquist. A partial's resonant frequency is set by its total round-trip time d_read + τ(ω), so upper partials, seeing a shorter loop, land sharp of the harmonic series — the stretched partials of a stiff piano string, exactly the physics that motivates the control.

Left there, the fundamental would sharpen too. The compensation is one line: m_d_read[v] = max(d_total − ap_tau, k_min_delay_samples), where ap_tau = allpass_phase_delay(m_ap_c, 2π·f/m_sr) is evaluated at the voice's fundamental. The main tap is shortened by precisely the phase delay the allpass adds at that frequency, so the fundamental's round trip is d_total again — pitch stays put while the overtones stretch. The max is the documented physical limit: at extreme warp × high tuning, d_total − τ falls below the interpolator's 2.5-sample floor, the loop cannot get shorter than the dispersion, and the pitch flattens — physical, and flagged in the maxref. The warp scenario pins the endpoints: warp 0 echoes at one period; warp 100 stays bounded, still resonates, and its tail correlates < 0.5 with the harmonic tail (a genuinely different spectrum, not a filter tilt).

phase: plucking the string at its midpoint

The output tap is out = y − pickup·read_linear(d_half) with d_half = d_total/2. Consider loop content at harmonic n of the voice — frequency n·f, i.e. n·fs/d. Delaying it by d/2 samples shifts its phase by

Δφ = 2π · (n·f/fs) · (d/2) = n·π

Even n: Δφ is a multiple of 2π, the delayed copy equals the original, and the subtraction (at pickup = 1) cancels it exactly. Odd n: Δφ = π, the copy is inverted, and subtraction doubles it. Pickup therefore sweeps continuously from the full series to odd-harmonics-only — plucking a string at its midpoint, where the even modes have a node. The test uses a 1 kHz voice at 48 kHz so the half tap lands on exactly 24 samples; Goertzel measures the 2f-to-f power ratio collapsing below 5% of its phase-0 value.

The parameter engine

All 22 parameters (k_num_params: gain, mix, three masters, warp, phase, freq/res/lp × 5) ride identical per-sample linear ramps. The bank keeps one count, m_ramps_active, and one flag, m_derived_dirty; process() advances only live ramps and recomputes the derived values (taps, fb, coefficients, mix gains) per sample while anything moves, once when everything has settled — the same two-tier idea as svf.h's coefficient cache, so steady state pays nothing for the smoothness.

The 16-slot preset morph is the same machinery pointed at all 22 targets at once: store_preset() snapshots the ramp targets (knob positions, not mid-ramp values); recall_preset(slot, seconds) calls ramp_to() on every parameter over n = seconds·m_sr samples. Because ramp_to() always retargets from r.current, a recall issued mid-morph — or a single slider grabbed mid-recall — is continuous by construction; no special case exists. The scenarios pin it: a 50 ms recall under a running sine produces no sample-to-sample jump above the click threshold and lands every parameter exactly on the preset; a mid-morph set_freq() reaches its own value while the other 21 keep morphing.

The engineering ledger

  • k_fb_max = 0.99999, applied after the rt60 solve. The cap makes "res 100 = longest possible resonance" a bounded statement; the rt60 math would happily request fb ≥ 1 for rt60 → ∞ (and res_master can push res_eff to 200).
  • Anti-denormal guard (|x| < 1e−15 → 0, the tap.comb~ idiom) on every recursive state — a comb ringing into silence otherwise decays into denormal territory and multiplies its CPU cost at the quietest moment.
  • Wet 1/5 normalization, a deliberate deviation. k_wet_norm = 0.2 scales the five-voice sum; the legacy abstraction wired five tap.comb~ objects straight into the output gain — a hot sum, +14 dB at full wet. The mix scenarios pin both endpoints: mix 0 is an exact passthrough, and mix 100 with all resonance off passes at unity to 1e−9.
  • Equal-power mix: m_dry_gain = cos θ · g, m_wet_gain = sin θ · g · k_wet_norm, θ = mix·π/200 — matching tap.crossfade~.
  • The pickup tap is linear-interpolated (read_linear), not Hermite — deliberate asymmetry: it is a feed-forward output tap, so its droop is applied once, not compounded per round trip; the argument that disqualified linear for d_read doesn't apply.
  • Allocation only in prepare(): one ceil(sr/k_freq_floor_hz)+8 buffer per voice (the 5 Hz floor honors the legacy 200 ms buffer). Every setter is a clamp plus a ramp retarget — safe while audio runs.

Checkpoint

A feedback comb is a string, and the file's one structural opinion is that the string's decay time — not its feedback coefficient — is the musical truth: fb = 10^(−3·d/(rt60·fs)), re-solved from the current delay so pitch sweeps preserve the ring. Hermite taps make the tuning real, the exact one-pole makes the damping knob honest, and the DC blocker retires the clipper by making high resonance a solved inequality (fine print stated) instead of a distortion stage. Warp is pure phase — dispersion paid back to the fundamental through allpass_phase_delay — and phase is pure geometry: half a loop is nπ, evens cancel, odds double. The rest is 22 ramps and one dirty flag, and the tests hold every claim.

Grains that sum to one: grm_pitchaccum.h

The user-facing chapter sold the effect on one image — +7 becomes +14 becomes +21, a staircase — and one engineering promise: "the tenth pass is as steady as the first." The image is a topology claim and the promise is an identity about a pair of window functions, and both are provable. This appendix proves them with the file's own names, then walks the pitch follower's failure mode and the ledger. The kernel scenarios in tap.pitchaccum_tilde_test.cpp drive tap::tools::pitchaccum::accum_bank directly and pin every measured claim; the sibling tap.shift~ tests pin the envelope identity to nine decimal places.

Transposition is a moving tap

A delay tap that moves changes pitch. Write the read position of a tap with delay D(n) samples behind a write head at sample n:

p(n) = n − D(n)

The output at sample n reproduces the input's phase at time p(n), so the output advances through the input at the rate:

dp/dn = 1 − dD/dn

A fixed tap (dD/dn = 0) plays at unity. A tap whose delay shrinks by (ratio − 1) samples per sample plays the buffer at ratio times real speed — up an octave means eating the delay line at one extra sample per sample. The code implements exactly this, inverted into phasor form: the tap's delay is base + window_samples · ph, and the phasor steps

m_phase += −(ratio − 1.0) / window_samples

so dD/dn = window_samples · dph/dn = −(ratio − 1), giving dp/dn = ratio. (The tt_shift provenance of that line is flagged in the code.) The transposition scenarios pin the result at the endpoints and the middle: +12 puts the energy of a 440 Hz sine at 880 Hz, −12 at 220 Hz, and 0 passes 440 untouched — each dominating its reference bin by the test's margins.

One moving tap cannot run forever — the phasor wraps, and at the wrap the tap teleports across the window: a splice. Hence the classic two-tap engine: a second tap rides the same phasor at ph_b = ph_a + 0.5 (mod 1), half a cycle apart, so one tap is always mid-window while the other is wrapping, and each is faded by envelope() so the splice happens at zero gain.

The envelope pair: an exact partition of unity

Here is envelope(ph, flank), region by region (flank ∈ (0, 0.5] is the crossfade width as a fraction of the cycle — m_flank maps xfade 1–100% onto (0.005, 0.5]):

ph ∈ [0, flank]:          sin²( π·ph / (2·flank) )      — cos²-shaped rise
ph ∈ (flank, 0.5]:        1                             — plateau
ph ∈ (0.5, 0.5+flank]:    cos²( π·(ph−0.5) / (2·flank) )— fall
ph ∈ (0.5+flank, 1):      0

The claim — stated as a comment in the file, and load-bearing — is that the taps at ph and ph + 0.5 sum to 1 exactly, at every phase and every flank width. Proof by the same four regions, writing e(·) for the envelope and using ph_b = ph_a + 0.5 mod 1:

ph_a ∈ [0, flank]:        ph_b ∈ [0.5, 0.5+flank]
    e_a + e_b = sin²(π·ph_a/2·flank) + cos²(π·ph_a/2·flank) = 1
ph_a ∈ (flank, 0.5]:      ph_b ∈ (0.5+flank, 1]
    e_a + e_b = 1 + 0 = 1
ph_a ∈ (0.5, 0.5+flank]:  ph_b wraps to ph_a − 0.5 ∈ (0, flank]
    e_a + e_b = cos²(π·(ph_a−0.5)/2·flank) + sin²(π·(ph_a−0.5)/2·flank) = 1
ph_a ∈ (0.5+flank, 1):    ph_b = ph_a − 0.5 ∈ (flank, 0.5]
    e_a + e_b = 0 + 1 = 1

Every crossfade region pairs a sin² with the cos² of the same argument; everywhere else a plateau pairs with a zero. The construction also joins each flank to its plateau with zero slope (the derivative of sin² vanishes at both ends), so there is no corner to click, and at flank = 0.5 the plateau vanishes and the pair degenerates to a complementary Hann pair.

Why demand exact? In a one-shot shifter, an envelope pair that sums to 1 ± δ is a gain ripple of δ at the grain rate — a subtle tremolo, mildly regrettable. But this object's entire identity is that the transposer sits inside a feedback loop: the ripple multiplies onto the signal on every pass, and after k trips the peaks have compounded to (1+δ)ᵏ while the troughs have decayed — a pumping that grows with exactly the feedback settings the effect is played at. That is why the original engine's window was replaced: tt_shift used a fixed 256-point padded-Welch table whose pair did not sum flat (the deviation is documented in both this file's header and tap.shift~'s). The identity is pinned numerically in the tap.shift~ wrapper tests — same envelope construction at flank = 0.5 — where DC at 0.5 pushed through moving taps at ratio 1.3 comes out equal to the input with max error below 1e−9: no grain-rate ripple, to double precision.

The accumulation topology

transposer::process() wires the loop in this order:

in ──►(+)──► delay buffer ──► two moving enveloped taps ──► y (out)
      ▲                                                     │
      └── × fb ◄── DC blocker ◄── m_fb_state (previous y) ◄─┘

The fed-back sample re-enters upstream of the taps: it is written into the buffer and then read back by the moving, windowed taps — which is to say it is delayed by delay_samples and transposed by ratio again. Every trip multiplies the frequency by another factor of ratio: +7 st becomes +14 becomes +21. Contrast the ordinary "feedback around a shifter" patch, where the feedback taps the delay output and re-enters the delay: each echo is re-delayed but shifted only once, and the staircase never climbs. The topology is the effect, and the kernel test pins its two-pass signature: a 440 Hz burst at +7 st, 300 ms delay, 70% feedback shows Goertzel energy at 659.26 Hz (one pass) in the 0.32–0.55 s window and at 987.77 Hz (two passes — the accumulation itself) in the 0.65–1.0 s window, each dominating the off-frequency reference bin.

The accumulation loop drawn against the ordinary shifter-in-feedback patch: upstream re-entry transposes every pass again; output-tap feedback shifts only once

Boundedness: the loop gain really is fb

The constant-sum envelope has a second payoff. Since e_a, e_b ≥ 0 and e_a + e_b = 1, the two-tap read is a convex combination of past loop samples — with linear taps its magnitude could never exceed the buffer's peak, and the Hermite taps can overshoot that bound only by the interpolator's small, bounded ripple. So the per-pass gain around the loop is genuinely fb — capped at k_fb_max = 0.99 — and not fb × (envelope ripple peak), which is the number that would have mattered with an uneven pair. The DC blocker in each loop kills the one component granular splicing can otherwise rectify into a ramp; its own tiny high-frequency shelf (2/(1+R) ≈ 1.0005) is absorbed a hundred times over by the 1% headroom in the cap — this kernel does not need the fine print that grm_comb.h's 0.99999 cap does. The test drives the worst case: both voices at 99% feedback, opposing transpositions, five seconds — output finite, peak < 50.

Modulation: one LFO, two seeded dice

Per sample, each voice's transposition is assembled as

trans_eff = m_ramp[p_trans1 + v].current + lfo + rnd
ratio     = 2^(trans_eff / 12)

The LFO is one global phasor (m_lfo_phase); voice 2 reads it at + modphase/360, so the two shadows breathe against each other at a settable phase — 90° by default. The random component is per-voice: tick_random() holds a target from a linear-congruential generator (seeded 1111 and 2222 in the constructor) and cosine-interpolates between held values at randrate, re-arming its phase when disabled so re-enabling starts a fresh segment rather than finishing a stale one. Deterministic seeding is a testability decision that is also a musical one — the same patch renders the same audio — and the test takes it literally: two identically configured banks render outputs with maxdiff == 0.0, bit identical. The modulation scenario pins the spectral effect from the other side: 1 st of 5 Hz LFO drains more than half the carrier bin's energy into sidebands.

The pitch follower, and the subharmonic trap

follow adapts the grain window to the input. The follower is deliberately cheap: input decimated by k_flw_decim = 8 (6 kHz at 48 kHz) into a 1024-float ring, and every k_flw_interval = 512 input samples a normalized autocorrelation over a k_flw_win = 512 window:

corr[lag] = Σ x[n]·x[n+lag] / √( Σ x[n]² · Σ x[n+lag]² )

searched over lags for 50–800 Hz. The naive readout — take the global maximum — has a classic failure mode that the code's comment names: a periodic signal correlates at every multiple of its period, so corr[2T] and corr[3T] sit at essentially the same height as corr[T], and windowing noise routinely pushes a multiple to the numerical top. The global argmax then reports a subharmonic — an octave or twelfth low — and the window snaps to twice the true period. The fix is two lines: find the global max best_r, then take the smallest lag whose correlation clears accept = max(k_flw_confidence, 0.85·best_r) — the earliest lag within 15% of the peak. Below k_flw_confidence = 0.6 nothing is accepted and period_s() reports 0: unpitched input is ignored rather than chased.

The window goal is k_flw_periods = 2 detected periods (clamped to the 5–200 ms window range), approached through a one-pole slew (m_window_eff_ms += 0.0005·(goal − eff), ≈ 40 ms time constant at 48 kHz) that relaxes back to the manual window value when follow is off or the gate says unpitched. Why two periods: at each grain wrap the tap jumps by exactly the window, so a window of k whole periods makes the splice displacement an integer number of cycles — the spliced waveform stays phase-coherent and the grain-rate artifacts land in tune with the source instead of at an arbitrary rate (the envelope cycle rate is |ratio − 1|·fs/window_samples, which at W = 2T scales with the source pitch, ≈ (ratio−1)·f₀/2). And k = 2 rather than 1 keeps the window above the 5 ms floor across most of the follower's 50–800 Hz range. The test pins both the adaptation and the trap: a 220 Hz tone converges the effective window into 6–13 ms, bracketing two periods (9.1 ms) — a subharmonic lock would demand 18.2 ms, outside the pin — and white noise leaves the window at the manual setting (70–100 ms around the 87 ms default).

The engineering ledger

  • 17 ramped parameters, same morph engine as grm_comb.h. One ramp array, m_ramps_active, m_derived_dirty recomputing derived values per sample while moving and once at settle; store_preset snapshots targets, recall_preset retargets every ramp from its current value, so mid-morph overrides are continuous with no special case. Pinned: an 80 ms recall under audio shows no sample-to-sample jump above 0.3 and lands every parameter to 1e−9; mix 0 is an exact passthrough (< 1e−9) even with 90% feedback churning inside the muted wet path.
  • The modulated ratio path lives in process(), not update_derived() — the code comments the split: LFO and random are inherently per-sample, so caching them would save nothing; the cacheable tier is delays, fb, voice gains, flank, and the equal-power mix.
  • Hermite taps with k_base_delay = 3 headroom — same 4-point interpolator and same ≥ 2-sample geometry constraint as grm_comb.h, here so that a voice delay of 0 ms is still legal under moving taps.
  • GRM's stereo-width fader is dropped, on purpose. The kernel is mono and the wrapper is single-channel by house rule (mc. wraps it); the omission is declared in the wrapper header and the maxref. Width would have been the only parameter that could not live in a mono kernel.
  • The follower is a mode, not a fader — set_follow(bool) sits outside the morphable parameter set (the file says so), because interpolating a boolean analysis mode over a morph is meaningless.
  • Allocation discipline. One buffer per voice sized for the worst case ((k_max_delay_ms + k_max_window_ms) at prepare-time sr, +16), a fixed 1024-float follower ring, and a stack-local correlation array; after prepare() the audio path allocates nothing, and the analysis cost — a few hundred multiplies per input sample, amortized — is paid only when follow is enabled.

Checkpoint

A tap whose delay changes at (1 − ratio) samples per sample is a transposer — dp/dn = ratio, by one derivative. Two taps on the same phasor half a cycle apart cover each other's splices, and the cos²/sin² envelope pair sums to one exactly, region by region, at any flank width — which in a feedback loop is the difference between loop gain fb and loop gain fb-times-ripple compounding every pass. The feedback re-enters upstream of the taps, so every echo is transposed again: the staircase is a topology, and the test hears both steps. The follower reads the earliest strong autocorrelation lag, not the tallest, because the tallest is routinely a subharmonic; two detected periods keep the splices phase-coherent. Everything else — ramps, morph, seeds, caps — is the same discipline as the comb bank, and equally pinned.

Two banks and a multiplier: vocoder.h

The user-facing chapter made three flat promises about tap.vocoder~: a silent carrier is silence, gain is exactly linear, and a silent modulator decays away at the follower rate. It could afford to, because none of those is a tuning outcome — each one is a structural fact about a very small graph. This appendix draws the graph, proves the facts, and then walks the three numerical choices (band placement, filter type, follower coefficient) that make the graph sound like a vocoder.

One honesty note up front. The original tap.vocoder~ source did not survive the revival; vocoder.h is reconstructed from the reference documentation — "a basic 24-band vocoder" with q and response_interval attributes. The topology below is the classic channel vocoder that documentation describes, and the tests pin its structural behavior; there is no lost binary to bit-compare against, and this chapter never pretends otherwise.

The graph: a bilinear form in 24 subbands

A channel vocoder is subband multiplication. Split both signals with the same filter bank, measure the modulator's level per band, scale each carrier band by that level, sum:

band i:   m_i = B_i(modulator)          the modulator through bandpass i
          env_i ← follower(|m_i|)        its envelope
          c_i = B_i(carrier)             the carrier through the identical bandpass
output:   y = gain · Σᵢ c_i · env_i

That is bank::process() verbatim — the loop body computes m, rect, m_env[i], c, and accumulates c * m_env[i], and the return line applies m_gain once to the sum. Three contracts follow from the shape alone, and tests/vocoder_test.cpp pins each one:

  • Silent carrier ⇒ exactly silence. Every summand carries a factor c_i. A biquad is linear with zero state at rest, so a zero carrier gives c_i ≡ 0 for all i, and the sum is identically zero no matter what the modulator (and hence the envelopes) does. The test drives a 220 Hz modulator against a zero carrier for 8000 samples and requires peak < 10⁻¹², but the true bound is exact: 0.0 * m_env[i] is 0.0.
  • Gain is exactly linear. The multiply c_i · env_i is the only nonlinearity in the graph, and it is bilinear — linear in the carrier with the envelopes held fixed, linear in the envelopes with the carrier held fixed. m_gain sits outside all of it, a scalar on the finished sum, and nothing upstream reads it. Two banks fed identical inputs with gains 1 and 2 must differ by exactly a factor of 2, float for float; the test requires |yb − 2·ya| < 10⁻¹² across 8000 samples.
  • Silent modulator ⇒ output decays at the follower rate. With the modulator silenced, rect = 0 and each envelope obeys env ← m_env_coef · env — a geometric decay with the follower's time constant. The output is bounded by Σ|c_i|·env_i, so it decays with the envelopes even while the carrier keeps playing. The test warms the bank up, silences the modulator for one second at 48 kHz (≈ 50 time constants at the 20 ms default — a decay of e⁻⁵⁰), and requires the late output under 10⁻⁴ of the warmed level.

The fourth pinned property, determinism (two identical runs compare equal with ==), is the repo-wide claim that the kernel is pure state-machine arithmetic: no randomness, no time, no allocation in the audio path.

The vocoder graph as a diagram: two identical filter banks, per-band envelope followers, 24 multipliers, and the summed gain

The bilinear form in 24 subbands — the graph shape the proofs read off.

Where the bands sit

Twenty-four bands span 50 Hz to 12 kHz, log-spaced. band_frequency(i) computes:

f_i = k_fmin · (k_fmax / k_fmin)^(i / (k_bands − 1))     i = 0 … 23
    = 50 · 240^(i/23)

so adjacent centres sit at a constant ratio of 240^(1/23) ≈ 1.269 — about 0.344 octave, a hair over four semitones, per band. Log spacing is the only defensible choice for this machine, twice over: the ear judges musical width by ratio, not by hertz, so equal-ratio bands devote equal perceptual width to each channel; and speech puts its identity (formants, the envelope the vocoder exists to capture) in the low kilohertz while its detail (fricatives) rides above — a linear spacing would waste twenty bands above 6 kHz and cram every vowel into two. The span itself brackets speech: 50 Hz is below any voice fundamental, 12 kHz is above any formant that matters, and recalc_filters() clamps each centre at 0.45 · m_sr so the top bands stay well below Nyquist at low sample rates rather than folding.

The filter: constant peak, unconditional stability

Each band is an RBJ Audio-EQ-Cookbook bandpass, the constant 0 dB-peak variant, computed in recalc_filters():

w0 = 2π · fc / sr        alpha = sin(w0) / (2·q)        a0 = 1 + alpha

b0 =  alpha / a0         a1 = (−2·cos w0) / a0
b1 =  0
b2 = −alpha / a0         a2 = (1 − alpha) / a0

"Constant 0 dB peak" is a normalization claim: the gain at the centre frequency is exactly 1, for any Q. It is worth proving, because the whole level architecture rests on it. Evaluate the transfer function at z = e^(jw0):

H(z) = alpha·(1 − z⁻²) / [(1 + alpha) − 2·cos w0 · z⁻¹ + (1 − alpha)·z⁻²]

denominator at z = e^(jw0):
  [1 − 2·cos w0 · e^(−jw0) + e^(−2jw0)] + alpha·(1 − e^(−2jw0))
  = e^(−jw0)·(e^(jw0) − 2·cos w0 + e^(−jw0)) + alpha·(1 − e^(−2jw0))
  = e^(−jw0)·(2·cos w0 − 2·cos w0) + alpha·(1 − e^(−2jw0))
  = alpha·(1 − e^(−2jw0))                    = the numerator exactly

so H(e^(jw0)) = 1 identically. Why it matters here: env_i is supposed to measure the signal's level in band i, and each carrier band is supposed to be scaled by that measurement and nothing else. With the constant-peak variant, changing q changes bandwidth only — the on-centre gain of all 48 filters stays pinned at unity, so the q knob narrows or overlaps the bands without re-balancing the reconstructed spectrum or re-calibrating the envelope levels. The cookbook's other bandpass (constant skirt gain) has peak gain Q; with the default q = 20 that would be +26 dB per band, scaling with the knob — every q move would also be a 24-band gain move.

Stability is likewise unconditional. A biquad is stable iff its coefficients sit in the stability triangle, |a2| < 1 and |a1| < 1 + a2. Here a2 = (1 − alpha)/(1 + alpha), which lies in (−1, 1) whenever alpha > 0 — and alpha = sin(w0)/(2q) is positive for any q > 0 and any 0 < fc < Nyquist; the second condition, 2|cos w0|/(1 + alpha) < 1 + (1 − alpha)/(1 + alpha) = 2/(1 + alpha), reduces to |cos w0| < 1, true on the same range. The code enforces the preconditions rather than assuming them: q is floored at 0.001 and fc clamped to 0.45·sr, so no attribute value and no sample rate can produce an unstable band. The sections run as Direct Form I (biquad::process keeps x1, x2, y1, y2) — at these moderate Qs and double precision, the plainest form is the honest one.

The follower: one coefficient, symmetric by construction

Each band's envelope is a one-pole lowpass over the full-wave rectified band signal:

rect     = |m_i|
env_i    ← m_env_coef · env_i + (1 − m_env_coef) · rect

with the coefficient computed in recalc_envelope() from the response_interval attribute:

tau        = response_ms / 1000                      (ms → seconds)
m_env_coef = exp(−1 / (tau · sr))

That is the exact one-sample step of a continuous first-order lag with time constant τ: the discrete pole e^(−T/τ) with T = 1/sr. So the documented "analysis period" is a time constant, precisely — after response_interval milliseconds of silence an envelope has decayed to 1/e of its value, and after a step up it has covered 1 − 1/e of the distance. Note what the code does not have: separate attack and release. One coefficient serves both directions, which is what the legacy surface documents (a single response_interval) and is why the user chapter calls the knob "the vocoder's attack and release." The 10⁻⁴ floor on response_ms keeps the exponent finite; at the 20 ms default and 48 kHz, m_env_coef ≈ 0.99896.

Why time-domain, when the siblings went spectral

tap.nr~ and tap.spectra~ (next chapter) are STFT machines. The vocoder deliberately is not, for three compounding reasons:

  • Zero algorithmic latency. The spectral scaffold costs exactly one FFT frame of delay by construction; this graph's output at sample t depends only on inputs up to t. A vocoder is played live against its carrier — latency is a musical defect here in a way it is not for noise reduction.
  • It is cheap. 48 biquads (5 multiplies + 4 adds each in DF I) plus 24 follower updates and 24 multiply-accumulates — on the order of three hundred flops per sample, no transform, no windowing, no frame buffers.
  • It is faithful. The original tap.vocoder~ was a real time-domain external; the pfft~-hosted abstraction that wrapped it in some patches only added smoothing and gain around it. Rebuilding it as an FFT effect would have been reconstructing a different object. So vocoder.h follows the svf.h/ladder.h idiom — prepare(samplerate) then per-sample process() — not the configure(fftsize) scaffold of the spectral set.

The engineering ledger

  • prepare() recomputes everything. It calls recalc_filters() (24 coefficient sets, each written into both m_mod[i] and m_car[i] — the banks are identical by construction, one computation assigned twice) and recalc_envelope(). set_q re-runs only the filters, set_response_ms only the envelope coefficient, set_gain is a bare store — each setter pays for exactly what it moves, the small-scale version of svf.h's two-tier update.
  • Setters are allocation-free and audio-safe. All state is in fixed std::arrays sized by k_bands; there is no allocation anywhere in the class, so the Min wrapper can forward attribute changes from the message thread while the perform loop runs.
  • The legacy surface is honored, with one documented fix. q and response_interval keep their documented names, meanings, and defaults (20 and 20 ms). The original registered both attributes as symbol; the wrapper (tap.vocoder_tilde.cpp) registers them as number, which is what they actually are — a Q value and a millisecond time — and says so in its header. gain is a small, admitted addition for level staging, since a band-multiplied signal lands quieter than either input.
  • Both banks clear together. clear() zeroes all 48 biquad states and the envelope array — the whole graph's memory, nothing else, so a clear message can never leave a stale envelope gating a fresh carrier.
  • What is deliberately absent: per-band gain trims, separate attack/release, a noise-driven "unvoiced" band — all classic vocoder extensions, all outside the documented surface being reconstructed. The reference page promised a basic 24-band vocoder; the file implements exactly that and stops.

Checkpoint

The vocoder is a bilinear form: two identical 24-band banks and one multiply per band. Everything the tests pin — silence in, silence out; exact gain linearity; follower-rate release — is a consequence of that shape, not of tuning. The numerics are three choices: log spacing (equal ratio per band, matched to hearing and to speech), the constant-peak RBJ bandpass (band level measures the signal, not the Q, and stability is a theorem with the clamps in place), and the exact one-pole coefficient e^(−1/(τ·sr)) (the documented period is an honest time constant, symmetric in both directions). Time-domain because latency, cost, and history all point the same way.

One STFT, three effects: fft.h, stft.h, nr.h, spectra.h

The user-facing chapters for tap.nr~ and tap.spectra~ both lean on the same claim: the machinery is transparent — set the effect to do nothing and the output is the input, exactly, one FFT frame late. All the trust in these objects lives in that claim, and it is not free: it has to be engineered into the window, the overlap, and one normalization constant. This appendix builds the stack bottom-up — the FFT, the scaffold, then the two small effects on top — and proves the transparency claim rather than asserting it.

fft.h: the transform, owned outright

The kernel repo's law is zero dependencies — plain C++17, standard library only. So the FFT is in-house: an in-place iterative radix-2 Cooley–Tukey in fft::transform(re, im, inverse), about forty lines. It is also one copy by design: the identical routine previously lived, byte for byte, inside conv_engine (tap.convolve~), tap.nr~, and tap.spectra~, and was consolidated at the kernel split so it is maintained and tested once — tests/fft_test.cpp is that single test point. The two halves:

  • Bit-reversal permutation. An iterative FFT consumes its input in bit-reversed index order; the first loop swaps each element i with its bit-reversed partner j, maintaining j incrementally (the carry-ripple idiom) rather than reversing bits per index; the i < j guard swaps each pair once.
  • Butterfly stages. For each length len = 2, 4, … N, combine pairs of half-blocks with twiddle factors e^(∓2πik/len). The twiddle is advanced by a complex-multiply recurrence (cwr, cwi rotated by (wr, wi)) — one cos/sin per stage instead of per butterfly. A recurrence accumulates rounding, but in double precision over these sizes it is far inside the pinned tolerance.

The scaling convention is asymmetric and load-bearing: forward is unscaled, inverse divides by N, so forward-then-inverse is the identity. Every claim is pinned in fft_test.cpp: the forward transform matches a naive O(N²) DFT to 10⁻⁹ for N ∈ {2, 4, 8, 16, 64, 256}; the round trip reconstructs random complex input to 10⁻⁹ at N = 128; a unit impulse transforms to an exactly flat unit spectrum; and a real cosine at bin 3 of 32 lands N/2 on bins 3 and 29 — fixing the sign convention (forward kernel e^(−i…)) and the conjugate-bin layout the effects below depend on.

stft.h: the scaffold, and the COLA proof

The STFT scaffold as a diagram: ring buffers and windows bracketing the FFT, the pluggable op, and the COLA-normalized overlap-add

One pump, two effects: nr and spectra are this pipeline with different middles.

stft is the overlap-add machinery shared verbatim by both effects: Hann window, fixed 4× overlap (m_hop = m_fftsize / m_overlap), a circular input buffer, a circular output accumulator, and a per-sample pump. process() takes the effect as a callable — op(re, im, N) mutates the N-point spectrum in place between the forward and inverse transforms; the only difference between tap.nr~ and tap.spectra~ is that lambda. The window, built in configure():

m_window[k] = 0.5 − 0.5·cos(2π·k / m_fftsize)        k = 0 … N−1

— the periodic Hann (denominator N, not N−1), which is what makes the overlap sums below exactly constant rather than rippling. The window is applied twice per frame: once at analysis (m_re[k] = inbuf·window[k]) and once at synthesis (outbuf += m_re[k]·m_window[k]·m_norm). With an identity op, the inverse transform returns the windowed frame exactly (the FFT round trip is the identity), so each input sample x is delivered to the output through every frame that covers it, weighted by w² each time:

y[t] ∝ x[t−N] · Σₘ w²(n − mH)          H = N/4, four frames cover each n

Perfect reconstruction therefore requires the shifted window-squared sum to be constant — the COLA (constant overlap-add) condition for double-windowing. For the periodic Hann at 4× overlap it is, and the constant has a closed form. Expand w²:

w²(θ) = (0.5 − 0.5·cos θ)² = 0.375 − 0.5·cos θ + 0.125·cos 2θ     θ = 2πn/N

A hop of N/4 advances θ by π/2. Over four hops, the cos θ terms are four quarter-turns of a phasor — they sum to zero; the cos 2θ terms advance by π per hop and cancel in adjacent pairs. What survives is the constant:

Σₘ w²(n − mH) = 4 × 0.375 = 3/2         for every n

The code does not hard-code 3/2. configure() overlap-adds overlap copies of m_window[k]² around a circular buffer and reads the value at cola[m_fftsize/2], setting m_norm = 1/c — for Hann at 4×, m_norm = 2/3 (verified numerically: the computed sum is 1.5 to within 10⁻¹⁵ at every index, so reading the midpoint is safe). Computing it keeps the scaffold correct for any window/overlap it might grow.

Latency is exactly N, and here is the accounting. The pump writes in[i] into m_inbuf[m_pos], reads the output from m_outbuf[m_pos], zeroes that slot, advances, and fires a frame every m_hop samples. When a frame fires, its index k holds input sample x[t₀ − (N−1) + k] (t₀ the newest sample, at k = N−1), and synthesis writes index k into m_outbuf[(m_pos + k) % N], which the pump reads k+1 samples later. Output time minus input time:

(t₀ + 1 + k) − (t₀ − N + 1 + k) = N          for every k, every frame

A frame cannot be transformed until it has filled — that is the whole cost, and why latency() simply returns m_fftsize. Both test suites pin the full contract at once: with a do-nothing effect (nr at threshold 0, spectra at remap 1), out[t] == in[t − N] to within 10⁻⁹ on broadband noise for all t ≥ 2N (the run-in covers frames that still window in zeros). This is the "transparent machinery" sentence in both user chapters, with its provenance attached: FFT round trip (pinned) × COLA constant (derived) × exact-N pipeline (derived).

nr.h: the gate, precisely

The spectral op is gate(), and its knee is short enough to quote in full as math. Per bin k:

mag  = √(re[k]² + im[k]²) · (2/N)
gain = 1                                if thr ≤ 0 or mag ≥ thr
gain = (mag / thr)^slope                if mag < thr   (slope ≤ 0 → 1)
re[k] *= gain;  im[k] *= gain

The 2/N puts mag on a sinusoid-amplitude scale (a real tone of amplitude A puts A·N/2 in each of its two conjugate bins; ×2/N recovers A). One honest calibration note: the frame is Hann-windowed before the FFT, and the Hann's coherent gain is 1/2 — a bin-centred sine of amplitude A actually measures mag = A/2 (verified numerically: A = 0.8 reads 0.400), with leakage in the adjacent bins. threshold is a linear amplitude on the windowed scale; a full-scale sine sits near 0.5, not 1.0.

The knee is a downward expander per bin. Take logs of the gain law below threshold:

L_out − L_thr = (1 + slope) · (L_in − L_thr)         in dB

Every dB below the threshold becomes (1 + slope) dB below it: slope 0 is unity (bypass by another name — the code special-cases it), the default slope 2 is a 1:3 expander, and slope → ∞ approaches a hard gate. Both re and im are scaled by the same real gain, so phase is untouched — the gate reshapes magnitude only.

Two structural notes. First, the loop runs over all N bins, mirror half included, with no symmetry bookkeeping — and needs none: the input frame is real, so its spectrum is Hermitian, magnitudes are symmetric (mag[N−k] = mag[k]), conjugate pairs get the same real gain, and Hermitian symmetry survives the op — the inverse stays real for free. (Hold that thought; spectra is not so lucky.) Second, the per-frame independence of the gain decision is exactly where musical noise comes from: a bin whose magnitude hovers near thr flips between pass and heavy attenuation frame by frame, and each isolated pass is one Hann-windowed near-sinusoid burst, milliseconds long — a chirp. Scattered over time and frequency, chirps sound like water. That is not a bug in the code; it is the knee's steepness meeting the frame rate, which is why the user chapter's cure is a gentler slope, not a different implementation.

tests/nr_test.cpp pins the three defining behaviors: gate open (threshold 0) reconstructs noise to 10⁻⁹ delayed one frame; a quiet bin-centred tone (amplitude 0.05 against threshold 0.5, slope 4) leaves a steady-state tail under 5 % of the input RMS; a loud tone (0.8 against 0.01) passes with RMS within 2 % of the input.

spectra.h: the remap, and why the mirror is not optional

The op builds a new spectrum over the lower half:

src      = lround(k · m_remap)                k = 0 … N/2
m_ore[k] = re[src], m_oim[k] = im[src]        if 0 ≤ src ≤ N/2, else 0

then forces DC and Nyquist real (m_oim[0] = m_oim[half] = 0), mirrors — m_ore[N−k] = m_ore[k], m_oim[N−k] = −m_oim[k] for 0 < k < half — and copies the scratch back over re/im.

The mirror is provable necessity, not tidiness. A real signal's DFT satisfies X[N−k] = conj(X[k]), and only Hermitian spectra invert to real signals. The remap fills the lower half by an arbitrary rule and touches nothing above Nyquist — the upper half still holds the input's bins, so the assembled spectrum is not Hermitian for any remap ≠ 1, and its inverse transform is genuinely complex. And the scaffold's synthesis reads m_re only — the imaginary part of the inverse is discarded. Keeping Re(IDFT(Y)) is algebraically inverse-transforming the Hermitian average ½(Y[k] + conj(Y[N−k])): without the mirror, the delivered effect would be an uncontrolled blend of the remapped lower half and the untouched upper half, half the intended signal shunted silently into the discarded imaginary part. The mirror makes the spectrum Hermitian by construction, so the inverse is exactly real and "keep the real part" loses nothing. DC and Nyquist are their own mirror images (k = N−k), so conjugate symmetry forces them real — hence the two explicit zeroes.

(The remap cannot run in place — output bin k may read a bin already overwritten — hence the m_ore/m_oim scratch, allocated in configure().) And the energy honesty: a remap is not a permutation, so Parseval is deliberately broken. For remap < 1, lround(k · remap) is non-strictly increasing — several output bins read the same input bin, duplicating its energy. For remap > 1 it strides — input bins are skipped, and every output bin above N/(2·remap) reads beyond Nyquist and is zeroed, discarding the input's top octaves. Neither direction conserves energy, and neither is meant to: the reference page has called this an "ultra-non-linear effect" since 2002; the kernel implements the rule, not a transform.

tests/spectra_test.cpp pins the two anchors: remap 1 reconstructs noise to 10⁻⁹ delayed one frame (the identity copies the lower half of an already-Hermitian spectrum, and the mirror rebuilds the upper half it started with); remap 2 moves a tone at input bin 16 to output bin 8 — lround(8 · 2) = 16 — verified by FFT-ing a frame-aligned slice of the steady-state output and requiring the peak at bin 8.

The engineering ledger

  • The effect is a template parameter, not a base class. stft::process takes SpectralOp&& and calls it once per hop; each effect passes a capturing lambda. No virtual dispatch in the audio path, full inlining.
  • Why the vocoder is not the third client. tap.vocoder~ is time-domain on purpose — zero latency, no frame, prepare(sr) instead of configure(fftsize) — see its own chapter. The spectral set accepts latency as a cost model; the vocoder's whole point is not paying it.
  • Allocation at configure() only. Window, in/out rings, FFT scratch, and (for spectra) the remap scratch are all sized there; process() is allocation-free. reset() flushes the running buffers without reallocating or touching the window — commented in the code as safe from a message handler, which is exactly how the wrappers use it.
  • fftsize is the one shared dial, and the scaffold makes its price explicit: resolution (bin spacing sr/N), smearing (a per-bin decision spreads over a whole frame), and latency (latency() returns N so the wrapper can report a true number to the host).
  • One FFT, tested once. The round-trip and DFT-reference pins in fft_test.cpp are what let this chapter treat "forward then inverse is the identity" as a premise everywhere above.

Checkpoint

The stack is three honest layers. The FFT is forty owned lines with an asymmetric scaling convention, pinned against a naive DFT. The scaffold windows twice, so reconstruction needs the shifted w² sum to be constant — for periodic Hann at 4× overlap it is exactly 3/2, measured rather than assumed, and the pipeline delays every sample by exactly N. On top, each effect is one spectral op: nr a per-bin 1:(1+slope) downward expander whose real gain preserves Hermitian symmetry for free; spectra a bin-index rule violent enough that reality — a real output — must be restored by explicit mirror. Transparency at the neutral setting is the theorem the whole stack exists to satisfy; both suites pin it at 10⁻⁹.

Seventeen, not four: diode_ladder.h

The transistor-ladder appendix derived why the Moog loop oscillates at k = 4. The 303's filter looks like the same idea — four capacitors, one feedback path — and behaves like a different species, because the diode ladder deletes the one luxury the Moog circuit has: buffering. Every diode pair both charges the next capacitor and loads the previous one. This appendix derives what that coupling does to the poles, why the oscillation threshold lands at exactly 17, why the shipping filter still refuses to self-oscillate at stock settings, and how the coupled nonlinear system is solved in closed form every sample.

Everything here is verified against Tim Stinchcombe's published TB-303 circuit analysis and the executed tb303.ipynb notebook, which matches the kernel's linearized response to his transfer function to 0.028 dB.

The chain: a diffusion line, not a cascade

With the diode conduction curve linearized (identity for now), the four node voltages obey a coupled chain:

v1' = ω·(S(u − v1) − S(v1 − v2))
v2' = ω·(S(v1 − v2) − S(v2 − v3))
v3' = ω·(S(v2 − v3) − S(v3 − v4))
v4' = 2ω·S(v3 − v4)          [the top capacitor is halved on the schematic]

Each middle equation has two terms — charge in from the left, charge stolen by the right. That is the loading, and it is the whole story: this is a discrete diffusion line, not four independent one-poles.

The diode ladder as four coupled capacitor nodes with charge flowing in from the left and stolen back by the right at every junction, a halved top capacitor, and the feedback path from the top node through the 150 Hz high-pass and the resonance gain back to the input sum

Every edge is a tanh; every middle node leaks both ways. The bidirectional arrows are what a buffered cascade doesn't have — and why the poles spread. Its normalized transfer function works out to exactly Stinchcombe's measured TB-303 response,

H(s) = 1 / (s⁴ + 6.727·s³ + 14.142·s² + 9.514·s + 1)

and the coefficients are not arbitrary — they are 4·2^(3/4), 10·√2, 8·2^(1/4): the equal-component chain with the top cap halved, which is also why Stinchcombe finds that changing that one cap shifts cutoff by 2^0.25. The poles are all real and spread ~25:1 (−0.13, −1.04, −2.33, −3.24 normalized). Compare the Moog ladder: four coincident poles. Consequences you can hear:

  • Asymptotically the slope is 24 dB/oct, but only ~14 dB falls in the first octave above cutoff — the honest version of the panel's "18 dB" claim.
  • At resonance 0 the −3 dB point sits ~3.2 octaves below the resonance frequency. The kernel's frequency parameter names the resonance peak, and the wide skirt below it is the real filter, not a tuning bug.

Seventeen: the closed loop's threshold

Feedback enters as u = drive·x − k·hp(v4). Ignore the high-pass for a moment and run Routh–Hurwitz on the closed loop: the stability boundary lands at exactly k = 17, with the marginal oscillation at √2× the stage rate. (Open303 normalizes its feedback by the same 1/17.) So resonance maps k = 17·resonance, putting 1.0 at the ideal chain's threshold — and the prewarp is chosen so that the √2 factor lands the oscillation on the labeled frequency:

g = tan(π·fc / fs_os) / √2       per stage (2g on the top stage)

Why a stock 303 never quite sings

Now put the high-pass back. The hardware's resonance feedback runs through a ~150 Hz one-pole high-pass (Open303's calibrated value, the fbhp default), and its phase lead pushes the would-be oscillation frequency up — to where the ladder attenuates more. Measured on the shipping kernel: the closed loop needs k ≈ 17.5 even at 8 kHz, ≈ 19 at 2 kHz, ≈ 25 at 500 Hz. The knob stops at 17. The emergent result — not programmed, derived — is the famous trait: a stock TB-303 never quite self-oscillates, and neither does this filter until you take the documented bend (resonance runs to 1.5, i.e. k = 25.5; past ~1.1 it sings at high cutoffs, slightly sharp of fc for the same phase-lead reason).

The high-pass buys two more behaviors for free:

  • Resonance thins as cutoff falls — low notes squelch instead of ringing, which is why a 303 keeps its bass at high resonance.
  • Closed-loop DC gain is exactly 1 regardless of resonance. The transistor ladder needed a comp parameter to buy its passband back; the 303's own circuit is the compensation, so this kernel has none.

Set fbhp 0 and the ideal analysis becomes exact: threshold at 1.0, oscillation at fc, drifting flat by ~0.7× the resonance excess past threshold — an amplitude effect (the growing swing saturates the edges unevenly), identical at every cutoff, and pinned by test.

The nonlinearity lives on the edges

In the circuit the coupling elements saturate — there are no buffer amps between stages to saturate instead. So the kernel puts tanh on every S(·) above: four saturators, one per diode-pair edge, slope 1 at the origin so small signals see exactly the linearized Stinchcombe response. There is no asym parameter here, deliberately: the diode pairs are complementary, so the transistor ladder's operating-point-mismatch story does not apply.

Solving the coupled system in closed form

The ZDF discretization (trapezoidal, as everywhere in the house) turns each sample into a system: five unknowns (v1..v4 and the loop input u) that all depend on each other through the couplings and the feedback. The kernel linearizes each edge with a secant gain γ = tanh(e)/e at an operating point, and then — this is the part worth reading in the code — eliminates the linear system bottom-up to closed form: v4 in terms of v3, then v3 = p30 + p32·v2, v2 = q20 + q21·v1, back-substituted until one division yields v1 and everything else follows. No matrix, no pivots, and unconditionally stable: every divisor is ≥ 1 for g > 0, γ ∈ (0, 1]. The feedback high-pass's state enters the same solve (its instantaneous gain 1 − G multiplies k), so the loop is closed exactly, high-pass included.

The two solvers differ only in how the secant gains chase the operating point:

  • solver_fast (default): solve at the previous sample's gains, refresh the gains at that solution, solve once more, commit.
  • solver_exact: repeat the refresh-and-solve until the node voltages move by less than 1e-12 (capped at 32 iterations).

Measured across a settings matrix out to resonance 1.4 and +24 dB drive — beyond hardware reach — the worst-case difference is −44.9 dBr, at 1.6–3.3× the CPU. The fast path's one correction is almost always enough because tanh is smooth and the operating point moves slowly at audio rate; the exact path exists so that claim never has to be taken on faith.

After the solve, the states advance trapezoidally using the true diode currents (tanh at the solved voltages, not the secant approximations) — the same "linearize to solve, commit with the real nonlinearity" pattern as ladder.h's one-pass commit.

Oversampling and the rest of the housekeeping

tanh generates harmonics; harmonics alias. The kernel runs 1×/2×/4× (default 2×) with zero-stuffing and matched 4th-order Butterworth anti-image/anti-alias cascades — the ladder.h pattern, self-contained here per the house rule against shared lookup tables. Every parameter rides a per-sample linear ramp; 16 preset slots morph through the same ramps; the right-inlet path recomputes the coefficient per sample for signal-rate cutoff. All state clears to zero, and all-zero state is a fixed point — a self-oscillating patch needs a ping, exactly like the transistor ladder.

The engineering ledger

  • Coupled solve vs. buffered shortcut. A "diode ladder" built from four buffered one-poles with new constants would miss the pole spread — the defining character. The 2×-larger algebra of the coupled solve is the price of the topology, paid once in closed form.
  • Secant linearization vs. Newton. Newton needs the derivative of five tanh terms through the elimination; secant gains reuse the same elimination unchanged and converge fast enough that solver_exact rarely iterates more than a few times. Same accuracy target, simpler code.
  • WDF: the documented no-go. A wave-digital rebuild of the same network was evaluated and declined (author-approved, 2026-07-18): solver_exact already converges the circuit's nonlinear equations, and a WDF would re-solve the same network differing only through the Shockley-vs-tanh diode curve, with no measured reference showing an audible delta to chase. The evidence lives in the notebook's solver A/B matrix.
  • No asym, no comp. Both absences are circuit facts, not omissions: complementary diode pairs, and a high-pass that is its own passband compensation.

Checkpoint

A diffusion chain whose transfer function matches the published analysis to 0.028 dB, with the oscillation threshold derived at k = 17 and then — because the feedback high-pass is modeled rather than idealized away — never reached at stock settings, exactly like the hardware. The nonlinearity sits on the coupling edges where the circuit puts it; the coupled ZDF system is eliminated to closed form and solved once or iterated to convergence, with −44.9 dBr between the two answers at settings the hardware can't reach. The character is the coupling, and the coupling is solved, not approximated away.

The couplings are the instrument: tb303_voice.h

The field-guide chapter argued that the 303 is unmistakable because its blocks are coupled — accent reaches the filter and the amplifier through shared circuitry with memory, slide is gate behavior, the envelopes have fixed interrelations. This appendix walks the per-sample code that implements those couplings: the measured envmod law, the C13 accent-sweep capacitor, the slide one-pole, the square shaper, and the phase-2 VCA. The filter itself is the previous appendix; this file composes it.

Sources, and the division of labor between them: Open303 (Robin Schmidt) supplies the measured constants — knob travels, envelope times, the envmod mapping, the square-shaper curve; the Devil Fish documentation (Robin Whittle) supplies the circuit behavior of the envelope/accent path, including the one place this kernel deliberately diverges from Open303. Every constant in the header carries its source.

One sample, in order

process() reads top to bottom as the signal path: pitch (with slide) → envelopes → the C13 update → the cutoff sum → oscillator and shaper → coupling high-pass → diode ladder → VCA → output coupling. Each stanza below is one of those steps.

Signal-flow diagram of the 303 voice with the couplings highlighted: the accent bus fanning to the envelope clock, the C13 charge path, and the VCA; the C13 capacitor feeding the cutoff sum

The file, as a schematic. Grey is what every clone has; red is what accent touches; amber is the cutoff CV that C13 leans on.

Slide: one coefficient, no special case

m_pitch += (m_pitch_target − m_pitch) · m_slide_coef

That is the entire slide implementation: a true RC lag (τ = the slide parameter, stock 60 ms — Open303's slideTime) on the pitch target. The gate logic makes it behave like the hardware: note_on with the gate low snaps m_pitch to the target before retriggering (a fresh note starts in tune); set_pitch with the gate held moves only the target, so the lag glides and neither envelope retriggers. Legato is slide — which is why the sequencer's gate-hold trick (see step_seq.h) needs no slide wire of its own.

Two envelopes, both RC discharges

The Main Envelope Generator and the VCA envelope are the same primitive — one-pole rise, exponential decay — with different constants and one coupling each:

  • MEG: 3 ms attack; decay = the decay knob (200 ms–2 s)… unless the note is accented, in which case the hardware bypasses the pot and runs at ~200 ms (accdecay, a bend, adjusts this clock). Faster and hotter is half of what "accent" means.
  • VCA env: fixed — ~3 ms attack (the Devil Fish "Soft Attack" bend widens it to 0.3–30 ms), a measured 1.23 s decay with no sustain, chopped by a 2 ms release at gate-off (Open303 measures ~1 ms; 2 is click-free). No knobs on the hardware, so no knobs here.

C13: the wow, as three lines of code

The accent sweep circuit is a diode feeding a capacitor through the resonance pot. The kernel's model is exactly that sentence:

drive = accent_knob · note_accent · meg
if (drive > c13)  c13 += (drive − c13) · charge     // diode conducts: τ = 47 ms  (47k·1µF)
c13 −= c13 · drain                                   // always draining: τ ≈ 150 ms

The diode gating (if drive > c13) is the memory: between closely spaced accents the drain doesn't finish, so the next accent starts from residual charge and peaks higher — the build-up. The notebook measures the cutoff peak growing ×1.94 across a run of accents and returning within ×0.998 once they stop. The cutoff contribution combines the capacitor voltage with a direct MEG term reduced by it (Devil Fish: "~100/147 of the MEG minus the capacitor voltage" — what rounds the first accent's curve):

res_mix = 0.3 + 0.7·min(resonance, 1)     // the pot is ganged with resonance
acc_oct = 2.0 · res_mix · (0.4·max(drive − c13, 0) + c13)

Two things to note honestly. The RC time constants are component-derived; the sweep span (2 octaves) and the 0.4 direct weight are informed approximations, flagged as such in the header. And this is the kernel's one deliberate divergence from Open303, which models its accent path as a plain 15 ms leaky integrator with no across-notes memory. The A/B was done for real — Open303 built and rendered side by side — and the Devil Fish circuit description won because the memory is documented hardware behavior. The divergence is recorded in the header, not buried.

The cutoff sum: a measured law, not a mixer

envmod is not "envelope amount into a summing node." Open303 measured the hardware's actual mapping (calculateEnvModScalerAndOffset), and the kernel uses those regression lines verbatim. With c the knob's log-position between the measured travel endpoints (302…2394 Hz):

scaler = (1−c)·(3.774·e + 0.737) + c·(4.195·e + 0.864)
offset = 0.0483·c + 0.2944
fc_eff = cutoff · 2^( scaler·(meg − offset) + acc_oct )

The offset term is the hardware's "gimmick": turning envmod up also injects a counteracting DC shift, so the sweep's resting point moves down as its depth grows — roughly 2/3 of the sweep lands above the knob position and 1/3 below. That interaction is why the knobs feel like a 303 rather than like a synth with the same ranges. Note acc_oct adds outside the envmod scaling: in the circuit the accent sweep injects directly into the cutoff sum, so accents quack even with envmod at zero.

The square that isn't

The 303's square is its saw pushed through a transistor shaper, and Open303 measured the resulting curve. The kernel takes the polyBLEP saw (vco.h's machinery), makes a half-cycle-shifted copy, and applies the measured shaper:

square = −tanh( 10^(36.9/20) · shifted + 4.37 )

That ~70× gain and the 4.37 bias produce the rounded, notched pulse whose spectrum audibly differs from an ideal 50 % square. The waveform parameter is a ramped blend between saw and shaped square, so switching glides click-free.

The couplings at the edges: two high-passes

Two one-pole high-passes bracket the filter — 44.5 Hz before it, 24.2 Hz after (both Open303-calibrated coupling corners). The post-filter one earns its keep twice: it is the output coupling, and in vca warm mode it absorbs the saturator's signal-dependent DC, which is exactly what the hardware's coupling capacitor does.

The phase-2 VCA: distortion that tracks the envelope

vca clean is a multiply — bit-identical to phase 1. vca warm models the one-transistor class-A stage as a slope-normalized biased saturator applied after the envelope gain and before the output coupling (the hardware order):

S(v) = ( tanh(d·v + b) − tanh(b) ) / ( d·sech²(b) ),   d = 2.0, b = 0.3

Unity slope at zero means quiet notes pass essentially clean; the bias means hot signals pick up even harmonics and compression. Because the envelope sits inside v, the distortion tracks it: measured 5.4 % difference-signal on quiet notes, 11.5 % on full accents, ~11 % second harmonic on a full-scale sine with ~−4 dB of compression. d and b are probe-calibrated informed constants — the header flags schematic-derived values as an audition-time refinement, which is the honest state of things.

Per-unit spread: seed/tolerance

The house vco.h convention, applied to a whole voice: tuning trim, cutoff scale, envelope times, slide and C13 RCs each take a deterministic per-seed offset scaled by tolerance, and the oscillator receives the seed plus a proportional imperfect amount. tolerance 0 is the nominal schematic, bit-identical to an unseeded voice (pinned by test); an mc. stack with different seeds drifts apart the way a wall of real units does.

The engineering ledger

  • One object, not a modular kit. The C13 path touches the MEG, the resonance knob, and the cutoff sum; accent touches the MEG clock, the VCA gain, and the sweep. Decomposed into osc + filter + env externals, every one of those wires would be the user's problem and most patches would omit them. The couplings live between the blocks, so the object boundary goes around them.
  • Measured constants over derived ones, where measurements exist. Open303's envmod law and shaper curve are adopted verbatim rather than re-derived from the schematic — they were measured against hardware, and re-derivation would add error, not rigor. Where Open303 simplifies (the accent memory), the circuit description wins instead. Each choice is sourced at the constant.
  • The wow's parameters are honest approximations. Sweep span and the direct weight await a hardware-calibration pass; the shape (diode gating, two RCs, resonance ganging) is circuit-derived and pinned by the ×1.94 measurement. Flagged, isolated, waiting — the autowah pattern.
  • process_at() per sample. Pitch (note + tuning + slide) can change every sample, so the oscillator is driven at signal rate rather than through a control-rate frequency parameter. The slide RC would be audibly steppy any other way.

Checkpoint

A voice whose per-sample loop is the schematic's block diagram: slide as one RC coefficient plus gate logic, envelopes as discharge curves with the hardware's fixed interrelations, accent as a hotter-and-faster MEG plus a diode-gated capacitor whose leftover charge is the wow, a cutoff law measured off real hardware complete with its gimmick, a square that is a shaped saw because that's what a 303's square is, and a VCA whose warmth tracks the envelope because the envelope sits inside the saturator. Every constant carries its source, and the one divergence from the reference implementation is documented with its reason.

One network, eight voices: the tr808_* headers

Roland built an entire drum machine out of about four circuit ideas, so the kernel does too: a bridged-T resonator class, a six-oscillator metal bank, a noise/VCA toolkit, and eight thin per-voice headers that compose them. This appendix covers the shared blocks' math — the bridged-T's trapezoidal solve and why the bass drum needs it solved that way, the metal bank's tolerance model — and the per-voice compositions, ending with the calibration pass that re-fit the family's envelopes against a real unit.

Provenance: the Werner–Abel–Smith papers (the DAFx-14 bass-drum analysis and the cymbal/cowbell companions) and the TR-808 Service Notes, read component by component. Every constant in these headers carries a schematic designator or a paper section; the calibration numbers live in tr808_calibration.ipynb.

bridged_t.h: the universal voice circuit

The bridged-T resonator drawn as a circuit: op-amp, capacitive arms, bridge and leg resistors, the exposed Vcomm node, and the kick's per-sample leg modulation

The network every voice reuses, with the kick's circuit-bending drawn in red.

An op-amp with a bridged-T network in its feedback path — capacitive arms C_a, C_b, a resistive bridge, a resistive leg to ground — rings when kicked, as a decaying pseudo-sinusoid at

fc = 1 / ( 2π · sqrt(R_leg_eff · R_bridge · C_a · C_b) )

where R_leg_eff is the leg in parallel with every resistive injection into the center node. Roland used this network in every voice: as the resonator of the kick, snare, toms/congas, rimshot, and claves, and as the band-pass of the clap, cowbell, cymbal, and hats. One class, one family.

Two implementation decisions matter:

  • The topology is reproduced, not summarized. With injections grounded, the class's transfer function matches the DAFx-14 paper's printed Eqn. (5) coefficient by coefficient (β₂ = α₂ = R_eff·R167·C41·C42, and so on); the injected paths match their Hbt2/Hbt3, interchanged by injection resistor; and the center node the paper calls Vcomm is exposed, because the bass drum's pitch-sigh nonlinearity reads it. The whole thing was re-derived by nodal analysis and pinned by unit test — the paper is trusted, then verified.
  • Trapezoidal on the states, not bilinear on the coefficients. The discretization uses capacitor companion models — a 2×2 linear solve per sample — which is algebraically the bilinear transform the paper uses, but solved on the network states directly. The reason is the bass drum: its leg resistance is modulated per sample (the attack shift shorts a resistor through Q43; the pitch sigh shrinks the effective leg through a fitted nonlinearity). With a coefficient-form biquad that would mean a full redesign every sample; with the companion-model solve, a time-varying resistor is just a changed matrix entry. Same ZDF family as the house svf.h.

The kick, since it exercises everything

tr808_kick.h composes the resonator with the paper's full block diagram: pulse shaper → retrigger network → bridged-T with a feedback buffer closing a regeneration loop → tone → level. The three signature behaviors are all emergent from the modeled schematic: for ~6 ms the envelope saturates Q43 and the ring sits near ~129 Hz (the attack punch); as the envelope collapses, C39/R161/D52 kick the center node again (the retrigger, so the note doesn't step down); and leakage lifts Q43's base when the center node swings below a diode drop — the paper's fitted memoryless nonlinearity (α = 14.315, V₀ = −0.556, m = 1.4765e-5) converts Vcomm to a collector current that shrinks the leg, so big early swings ring sharp and relax down as the note decays. That is the sigh, and it is a different mechanism from the attack jump — the paper's central untangling, preserved here. One erratum survives in the header: the paper's Eqn. (9) as printed is garbled, so the leg formula was re-derived from KCL at Q43's collector and matches their stated limits.

Accent is the trigger voltage — 4–14 V on the bus, mapped from the 0..1 edge amplitude — exciting the network harder, not scaling the output. And filter states persist across triggers, so rolls interfere with the ringing tail: no machine-gun effect, by construction rather than by crossfade.

swing_vca.h: the small shared parts

The 808 shapes its percussive gains with one-transistor "swing type" VCAs driven by RC discharges, not ADSRs. The header holds the three primitives the noise voices share: decay_env (one-pole rise to a level, exponential decay — retriggering re-aims the rise, no reset click), the linear swing_vca gain (the hardware's "many high harmonics" are a flagged refinement), and white_noise — a seeded xorshift64*, because the 808 has exactly one noise generator feeding the snare's snappy, the clap, the maracas, and the toms' noise layer, and because determinism-per-seed is a house invariant: renders reproduce, tests pin, mc. instances decorrelate.

metal_bank.h: six squares and a spread

The metallic voices all draw on one bank of six Schmitt-trigger relaxation oscillators: nominal 205.3, 369.6, 304.4, 522.7 Hz plus the two trimmer-tuned at 800 and 540 (the pair the cowbell taps), duty 47.98 % per the paper's HD14584 analysis. Three modeling calls:

  • Naive squares are faithful. The fundamentals sit below 1.2 kHz and the hash above them is immediately band-passed; the residual aliasing folds into the same inharmonic wash the circuit itself produces. PolyBLEP would be cost without benefit — a rare sentence in this repo, so it's documented.
  • Tolerance is part of the instrument. The RC parts put any given unit's oscillators up to ~20 % off nominal — the paper's measurement, and the reason no two 808s' cymbals sound alike. tolerance scales a deterministic per-seed spread of exactly that width. This is not "analog warmth" seasoning; it is a measured production statistic.
  • The two band-pass voicings (~3440 and ~7100 Hz, Q fit to the paper's published skirts) and the Q19 attack smoother (τ = 102.44 µs less a 0.7258 V base-emitter drop, their least-squares fit) live here too, because cymbal, hats, and cowbell all share them.

The voices, as compositions

Each tr808_*.h is a thin arrangement of the blocks above, with its own schematic constants:

  • Snare: two bridged-Ts at the late-revision ~173/336 Hz (the design change is documented in the header), a trigger divider, and the snappy path — decay_env-shaped noise, band-limited near 4 kHz.
  • Clap (clap|maracas): ~2 kHz dual band-pass noise through a VCA driven by the Service Notes' Figure-13 three-teeth sawtooth — the "multiple hands" transient — plus the Q70 reverberation tail.
  • Hats: one circuit, two envelope paths, and the hardware choke (Q23/R173): a closed-hat trigger terminates a sounding open hat, pinned by test. This is why tap.808.hat~ is one object with two inlets — the choke is unimplementable across separate externals.
  • Cymbal: the bank through both voicings with two separately enveloped bands (strike/ring/body), decay spanning the chart's 350–1200 ms.
  • Cowbell: just the 540/800 pair into the ~860 Hz voicing, two-slope envelope.
  • Toms/congas (@size × @model): the resonator at the chart tunings with the D80/D81 attack pitch fall; toms add a pink-noise layer (pinned by seed-sensitivity, since the diode bend's own harmonics defeat spectral separation); congas are the same circuit, no noise.
  • Rim/claves: the ~1667 + 455 Hz crack with the swing-VCA's tanh harmonics, versus the pure ~2500 Hz tick.

The family also carries per-channel summing gains (k_tomc_mix, k_cl_mix — the hardware's summing resistors into the mix bus): the bridged-T's impulse gain grows with fc·Q, and before the balance pass the high conga peaked at ~5.3 while other voices sat far lower. Every voice's full-accent peak now lands in a consistent ~0.3–1.0 band, pinned by test.

The calibration pass: what measurement actually changed

The family was calibrated against a real unit (s/n 103852) recorded from the individual outs with knob positions encoded in the filenames — a 0/2.5/5/7.5/10 dial grid, 116 samples — so the comparison ran per knob cell, with identical measurements (spectral-peak fundamental, −40 dB decay, power centroid) on both sides. The result is a clean split:

  • Frequencies: the schematics were right. Kick within 2.4 %, snare within 1.2 % (including the tone-max mode flip), toms/congas/cowbell/ claves within ~4 %. The kick needed no constant changed.
  • Time: the recordings won. Tom, conga, cowbell, and clap tails roughly doubled; the snappy was band-limited and re-enveloped; the rimshot re-voiced low-dominant; the cymbal's decay span and brightness corrected; the closed hat's brightness residual later resolved by the hats' sizzle blend.

Each header carries its calibration note with numbers and residuals. The lesson is worth stating as a rule: schematics get you the frequencies; recordings get you the envelopes — decay behavior hides in pot tapers, electrolytic tolerances, and aging that no schematic states.

The engineering ledger

  • One resonator class vs. per-voice filters. Eight voices reduce to ~4 blocks plus thin compositions only because the bridged-T class keeps the injected-path structure of the real network instead of collapsing to a generic biquad. The generality was free once the nodal analysis was done — and the kick's per-sample leg modulation required it.
  • Behavioral envelope generators. The kick's EG is modeled as fast rise / ~1.1 ms release rather than as its own transistor network — the paper's own simplification, adopted with its citation. Fidelity effort went where the analysis said it matters (the leg, the retrigger, the sigh), not uniformly everywhere.
  • The WDF door, left closed but unlocked. The flagged @circuit upgrade path (wave digital, the svf.h two-circuit pattern) remains gated on an A/B showing an audible delta the informed model misses. The DAFx-14 paper's own finding — device nonlinearity matters less than folklore claims — suggests the gate may never open, which would itself be a documented result.
  • Determinism everywhere. Seeded noise and seeded tolerance mean every render, test, and calibration measurement is reproducible bit-for-bit. The calibration pass would have been guesswork without it.

Checkpoint

One network class matching the published transfer functions exactly and solved on its states so a time-varying resistor costs nothing; a metal bank whose ±20 % spread is a measurement, not a vibe; voices that are thin compositions with schematic-designated constants; and a per-knob-cell calibration pass that confirmed the frequencies, corrected the envelopes, and wrote its residuals into the headers it changed. Four ideas, eight voices, every number traceable.

Time as a function of phase: step_seq.h

The sequencer header is the smallest DSP file in the kernel and the one whose central decision does the most work per line: the engine owns no clock. It is handed a phase — a number in [0, 1) meaning "here is where we are in the pattern" — and everything else (the current step, whether this sample is a boundary, how far through the step we are) is derived from it, statelessly, every sample. This appendix explains why that one decision buys sample accuracy, polymeter, scrubbing, and drift-free multi-row lock for free, and then walks the three pieces built on it: the swing warp, the two emitters, and quantized recall.

Verification lives in two places: tests/step_seq_test.cpp (19 Catch2 scenarios, including a pairing test against the real tb303_voice.h) and the executed step_seq.ipynb. The design of record is plans/tap.seq.md in the Max package repo.

Deriving the step, in O(1)

Ignore swing for a moment and the whole clock is one line:

k = floor( wrap(phase) · length )

Swing delays each odd-numbered step's start by swing/2 of a step, so the start of step k is

start(k) = ( k + (k odd ? swing/2 : 0) ) / length

and the derivation gains one correction: compute the naive k, and if it is odd but the fractional position hasn't yet reached swing/2, the sample still belongs to the (even) step before it. Two comparisons, no search — the boundaries are monotone, so the correction is exact.

A step entry is simply k != k_previous. That definition, rather than "the clock ticked," is what makes the engine indifferent to how the phase moves: run it backwards and entries still fire (pinned by test); jump it and the landing step fires once; feed it a constant and nothing happens after the first sample. reset() just forgets k_previous, so a transport start fires its downbeat.

Time as a function of phase, drawn: a phase ramp through the floor derivation and step-entry inequality into the two emitters, plus two row lengths reading one ramp without drift

The one decision the file turns on — and polymeter falling out of it as arithmetic.

Why phase, not a pulse clock

The alternative — count incoming clock pulses — is how most step sequencers are built, and every one of them then grows a reset input, a position protocol, and a drift story. Deriving from phase dissolves all three:

  • Sample accuracy is inherited from the phase source. The notebook measures trigger edges landing within one sample of the analytically computed boundaries — the one sample being float rounding at the boundary itself, not accumulated error.
  • Multi-row lock is structural. Two rows fed the same ramp cannot drift, because neither owns any timing state that could drift. Mute one for an hour; it re-enters in place.
  • Polymeter is arithmetic. A length 12 row against length 16 rows off one ramp divides the same cycle differently — 12 and 16 entries per cycle, measured. The TR-808's triplet "pre-scale" falls out as a special case.
  • Position is explicit. Scrubbing, reversing, and jumping are the caller's choices about the ramp, not features the engine implements.

The cost is honest too: the engine cannot free-run. That is deliberate — phasor~ (transport-locked or not) already exists, and a sequencer that owns tempo is a sequencer that fights the transport.

Position within the step, and the gate duty

The tick also reports pos — the fraction of the current step's actual (swung) span elapsed — computed from the same start() function. Gate timing hangs off it: the note row closes its gate at pos ≥ 0.5, the pinned Open303 duty. Measuring duty against the swung span rather than the nominal step means gates never collide however hard the swing is pushed.

The trigger row: an impulse and a re-arming gap

trigger_row is the small emitter: on entry to a sounding step, emit the step's velocity for one sample (or pulse_ms worth, for envelope consumers), else zero. The single-sample default is a contract, not a simplification: every downstream tap.808.* voice re-arms its edge detector below 1e-3, and the test suite pins that two adjacent sounding steps produce two clean detectable edges. The header documents the one way to defeat this — a pulse_ms longer than a step merges back-to-back triggers — rather than silently preventing it.

The note row: a five-state sentence

note_row implements the tap.303~ contract, and its entire behavior fits in one paragraph of code. On entering step k: if the step is gated and its slide flag is set and a note is already sounding, change the pitch output and leave the gate level alone — that is legato, and the voice's RC does the glide. If gated without that condition, set the gate to 1.0 (2.0 if accented) — a fresh edge. If not gated, drop the gate. Between entries: close the gate at the duty point unless the next step is gated and slid — that look-ahead read is the gate-hold, and it is read live from the pattern each sample so an edit lands immediately.

Three edge cases are worth naming because the tests pin them:

  • Slide from a rest is a plain trigger — there is nothing sounding to slide from, so the flag degrades gracefully (the voice's note message behaves identically).
  • Chained slides chain — each held boundary defers the duty close to the next step, so a run of slid steps is one unbroken gate. Sixteen gated steps with three slide flags produce exactly thirteen note-ons, measured.
  • The wrap is a boundary like any other — a slide from step 15 into step 0 holds across phase 1→0, because nothing in the derivation treats the wrap specially.

One convention deserves its provenance note: the slide flag sits on the target step (the note being slid into), matching the package's note <pitch> [accent] [slide] message and the original interface dry-run. The hardware stores the flag on the source note ("slide to next"). The data models convert trivially — shift the flag column by one — and the divergence is documented in the header rather than discovered by a user.

Quantized recall: swap on the boundary sample

Patterns live in 16 slots. recall arms rather than acts (unless quantize now): the armed slot is applied on the next cycle entry (step 0) or step entry, and — the detail that keeps it exact — the engine then re-derives the current step against the new pattern's grid on that same sample, since the new pattern may have a different length. The notebook pins the semantics end to end: armed mid-cycle, the running pattern finishes its bar at its own amplitudes, and the first trigger after the wrap carries the recalled pattern's. That one message is the TR-808's A/B-half and basic/fill switching.

What is deliberately absent

No randomness (bit-exact by construction, still pinned by test, because invariants that aren't tested rot). No allocation after prepare() — the pattern store is a fixed 64-step array times 16 slots. No run/stop, no direction modes, no ratchets: the first two belong to the phase source, and the last is a future emitter, which is the point of the next paragraph.

The engineering ledger

  • Engine/emitter split. The clock math lives once; trigger_row and note_row are each a screenful. A future row flavor — CV, probability, ratchet — is another emitter, not another clock. This is also why the Max-side question "one generic object or two family objects?" could be answered by product taste rather than by implementation cost.
  • Look-ahead vs. cached hold. The gate-hold could cache "next step slides" at entry; reading it live costs one array access per sample and makes pattern edits take effect mid-step. Cheap beats stale.
  • Sample-resolution boundaries. Sub-sample trigger placement (fractional edge amplitudes à la BLEP) was considered and declined: the consuming voices detect edges at sample resolution, so sub-sample machinery would add complexity no consumer can observe. If a future voice interpolates its trigger time, the tick already carries the information needed to add it.
  • The armed-recall re-derivation. The subtle bug in naive quantized recall is applying the swap after deriving the step, leaving one sample computed against the old grid. Applying, then re-deriving within the same call, is two extra lines and the difference between "exact on the wrap sample" (measured) and "usually fine."

Checkpoint

A sequencer that is a pure function of phase plus a pattern: one line of derivation, one comparison for swing, entry as inequality — and from that, sample accuracy, polymeter, reversibility, and drift-free lock without a clock to maintain. The rows translate steps into the two shipped voice contracts, with slide as a held gate and a live look-ahead; recall swaps patterns on the exact boundary sample. Nineteen scenarios and an executed notebook agree, and the most satisfying number in either is small: thirteen note-ons, for sixteen steps, three of which arrived without knocking.

Three ways to move a pitch: yin.h, psola.h, pvoc.h

The user-facing chapter promised that three interchangeable engines land the same intonation. This appendix is about why that is hard: each engine is a claim about what a pitched sound is, and each claim fails somewhere specific and measurable. Two of those failures were found the good way — as failing tests during development — and both are now pinned as contracts rather than patched into vagueness.

These three headers live in the shared DspTap repository (the same home as the real FFT that machine/spectral.md describes), because a pitch detector and two shifters are not Max material or even TapTools material — they are primitives, in the fft.h mold: a double-precision golden model, a float32 embedded profile pinned against it, allocation-free noexcept processing, fixed documented latency, and hot loops kept contiguous as future Helium/HVX backend seams. Every number below is produced by the shipping code through DspTap's C ABI in notebooks/pitchshift.ipynb, and gated in test_yin.cpp / test_psola.cpp / test_pvoc.cpp.

A period is a lag that explains the signal: yin.h

Autocorrelation says: a signal is periodic at the lag where it best matches itself. The trouble is that a harmonic-rich signal matches itself rather well at twice the true period too, and "rather well" wins often enough to make naive autocorrelation an octave gambler. YIN (de Cheveigné & Kawahara, 2002) replaces "best match" with "smallest failure": a squared difference function

d(τ) = Σ (x[j] − x[j+τ])²,   j over the integration window

then divides each lag's failure by the running mean of all failures up to that lag — the cumulative-mean normalization — so d′(0) ≡ 1 and small lags stop being free wins. The detector takes the first lag whose normalized failure dips under an absolute threshold (0.1 by default), descends to the local minimum, and refines it with a parabolic fit over the three surrounding values — the sub-sample step that turns an integer lag grid into a fractional period.

The contract, measured: worst sine error 0.17 cents across 82–988 Hz (including deliberately non-integer periods), worst sawtooth error 0.155 cents with no octave errors — the trap the normalization exists to disarm. Noise and silence report unvoiced, and the threshold gates honestly (a deliberately dirtied sine flips to unvoiced when the threshold is tightened below its measured aperiodicity).

One honest limit, kept on purpose: first dip under threshold scans from short lags to long, so on synthetic material whose fundamental is nearly absent — a formant bump with almost no energy at f₀ — a subharmonic lag that happens to land on an exact integer can dip deeper than the true period's slightly-off-grid dip, and the detector follows it. The notebook demonstrates this deliberately and measures such material with a cepstral oracle instead. Real voices keep enough fundamental that the rule holds; the failure is documented, not hidden.

Finding 1 — a shifter that moves everything except the envelope: psola.h

TD-PSOLA's move is disarmingly physical. Put an analysis mark every period. Cut a two-period Hann grain around each mark. To synthesize a new pitch, lay the grains back down at a new spacing — period/ratio — and sum. The windows are arranged to sum to one at the identity, grains are scaled by 1/ratio to keep the sum flat elsewhere, and synthesis marks are placed with sub-sample precision (each grain resampled through the same 4-point Hermite kernel the rest of the family uses) so mark rounding never becomes pitch jitter. The file adds one real-time honesty: marks come from a free-running period-synchronous scheduler, not glottal-epoch estimation — the standard practical simplification — and the caller supplies the period, so detector and shifter stay independently testable.

Then the finding, told as it happened. The first shift-accuracy test fed the shifter a pure sine at ratio 2 and got back 0.0000 — silence, from a correct implementation. Because that is what PSOLA does: re-spacing period-synchronous grains resamples the source's spectral envelope at the new harmonic spacing. A voice's envelope is wide — formants — so the new harmonics sample it fine, which is exactly the celebrated property: formants stay put while pitch moves. A pure sine's envelope is a single spike at f₀, and after an octave up the new harmonic grid (2f₀, 4f₀, …) contains nothing at f₀ — the output honestly, correctly vanishes. One property, two faces.

The response was not to patch the algorithm into something less itself. The shift tests were rewritten onto voice-like material (a normalized band-limited sawtooth: ±8 cents across ratios 0.5–2.0 at healthy level), and the pure-tone behavior got its own pinning test, PureToneOctaveUpThinsOut, so that if this property ever changes, someone is forced to explain why. The header now opens with the warning label: know what PSOLA is; feed it harmonics; for pure tones use a waveform-preserving shifter.

Latency is fixed at 2·max_period + 2 samples — the price of grains that must be fully received before they can be laid back down.

Finding 2 — the naive phase vocoder loses half its level: pvoc.h

The textbook pitch shifter looks like four honest lines: STFT with Hann windows at 4× overlap; per-bin instantaneous frequency from the frame-to-frame phase increment; remap each analysis bin k to synthesis bin round(k·ratio); accumulate each synthesis bin's phase at its scaled frequency and inverse-transform. It is in tutorials everywhere. Measured on a unit sine, it delivers 0.14–0.46 of the input level at fractional ratios — more than half the signal simply gone — and its "identity" at ratio 1 is a sine of the right frequency with the wrong waveform.

Two structural reasons. First, a single partial does not live in one bin; it lives in a Hann mainlobe pattern across four-ish bins, and round(k·ratio) scatters that pattern (220 Hz × 1.5 lands on bins {5, 6, 8, 9} — nothing at the true target, 7.04). Second, free-running per-bin phase accumulators destroy the phase relationships across the lobe, and the overlap-add — which is a resampling filter with real opinions — partially cancels what remains.

The shipping design is Laroche–Dolson peak-region shifting. Find the spectral peaks (local maxima over ±2 bins, gated 80 dB below the frame's strongest bin so the noise floor cannot claim regions). Split the spectrum into regions around them. Translate each region rigidly by an integer bin offset — the lobe pattern survives intact, phase relationships and all — and rotate the whole region by a single accumulated per-hop phase ψ.

And here is the bug that cost an afternoon and earned its own comment block: ψ must accumulate the full per-hop frequency difference, ψ += 2π·hop·f·(r−1)/N, not the sub-bin residual left after the integer shift. An integer bin shift is implemented, in effect, by a modulator e^(2πi·shift·n/N) — but n is the frame-relative sample index, so that modulator restarts every frame and contributes nothing to frame-to-frame phase advance. Subtract the shift from ψ (the "obvious" refinement) and every frame disagrees with the last about where the shifted partial's phase should be; the overlap-add quietly shreds the signal. With ψ carrying the full difference, the measured contracts land: sub-cent frequency accuracy at every tested ratio, ~0.95 level everywhere, and — because at ratio 1 every shift and every ψ increment is exactly zero and the analysis phases pass straight through — exact waveform identity, one frame late, to 7.8 × 10⁻¹⁶.

The envelope as a filter: LPC formant preservation

set_formant(true) adds the classic source-filter correction. Per analysis frame: autocorrelate the windowed time frame to lag 48, run Levinson–Durbin (always in double — an order-48 recursion in float32 is not a place to economize), and evaluate the prediction polynomial's magnitude over all bins with one extra FFT of its coefficients — the same transform engine, one more call. That gives a spectral envelope E(k) = 1/|A(e^jωk)|, and every relocated bin trades envelopes: content moving from bin k to bin j is scaled by E(j)/E(k) (clamped to ±24 dB so a near-zero envelope cannot mint gain). The excitation moves; the envelope stays. Measured: a synthetic 800 Hz formant on a shifted-up-a-fifth voice stays at 800 Hz with the flag on (band-energy ratio 62:1) and dutifully chipmunks to 1200 Hz with it off. At ratio 1 the correction is E(k)/E(k) — exactly unity — so the identity contract survives the feature untouched. The method is implemented from the published literature only, which in this corner of DSP is a policy statement, not just a citation habit.

The house pattern

All three files repeat the fft.h discipline because it keeps paying: basic_*<Sample> templates with double as the golden model and float as the embedded profile, cross-precision agreement pinned by tests; geometry fixed at construction and every buffer allocated there; noexcept, allocation-free processing; latency as a number in the header, not a vibe; and the expensive inner loops (YIN's difference function above all) written as plain contiguous arithmetic so a Helium or HVX backend can slot in behind the same contract with the scalar build remaining the oracle.

Checkpoint

A detector that measures failure-to-match instead of match, normalized so short lags stop cheating, refined below the sample grid — sub-cent, octave- safe, honest about the one synthetic that fools its first-dip rule. A grain shifter whose deepest property — resampling the spectral envelope — is both its celebrated feature and its pure-tone failure, pinned from both faces. A phase vocoder that works because peaks move as rigid families with one phase register each, carrying the full frequency difference — since the integer shift's modulator restarts with every frame. And an LPC envelope trade that lets the excitation move while the mouth stays. Three claims about what a pitched sound is; three sets of receipts.

The nearest allowed note: tune.h

The pitch-primitives appendix built the parts: a detector and two shifters, each with a numeric contract. This appendix is about the composition — tap::tools::tune::corrector, the object behind tap.tune~ — where the interesting problems are not algorithms but policies: what runs when, what is allowed to allocate, what happens when the detector reports nothing, and one measured surprise that became the kernel's best war story. The scenarios in tune_test.cpp drive the class directly, using the DspTap detector as an independent pitch oracle on the output; notebooks/tune.ipynb re-measures the headline claims through the C ABI.

The pipeline and its clock

process() runs per sample; analysis runs per hop (256 samples at 48 kHz, about 5.3 ms, scaled with the rate). Each sample: feed the detector's input ring, maybe analyze, advance two slews (the applied correction and the grain window), compute the ratio, resynthesize.

The geometry trick that keeps every setter real-time safe: the detector is built at prepare() for the worst case — lags from 2 kHz down to 55 Hz — and the user's set_range() merely filters results afterward, treating out-of-range estimates as unvoiced. Changing the range never reallocates, so it is safe mid-audio, and the price is a fixed analysis cost: with window = τ_max = 873 samples at 48 kHz, the YIN difference function is roughly 760k multiply-adds per analysis — an ~80 µs scalar spike every 5.3 ms, well inside a 64-sample vector's 1.3 ms budget, and the FFT-accelerated difference function remains available behind the same contract if an embedded target ever objects.

Analysis converts the detected period to MIDI, chooses a target (next section), sets the correction goal in semitones, and retargets the grain window. Unpitched frames set the correction goal to zero — the corrector relaxes toward honesty — while the window holds its last value rather than lurching toward a default.

tap.tune~'s pipeline as a diagram: the per-hop YIN/target/glide/ratio brain over the per-sample ring and resynthesis backend seam

The pipeline and its two clocks — the dashed region runs per hop, everything else per sample.

The period lock, told as a bug hunt

The first version of the resynthesis stage was the two-tap tap.shift~ engine with its grain window clamped to a sensible fixed range, minimum 5 ms. The oracle tests immediately failed — not wildly, musically: a 452 Hz input hard-snapped to A440 came out at 441.4 Hz, 5.4 cents sharp. Detection was exonerated first (0.06 cents), then the applied correction (right to five decimals). The bias lived in the shifter itself, and an isolation experiment found the shape of it:

grain windowmeasured outputerror
exactly 2 detected periods (212.4 smp)439.98 Hz−0.06 cents
fixed clamp (240 smp)441.38 Hz+5.41 cents
480 smp (≈4.52 periods)441.36 Hz+5.34 cents

The two taps ride the same phasor half a cycle apart, so they sit window/2 samples apart in the delay line. When window/2 is an integer number of source periods, the taps read the same phase of the waveform and their crossfade is invisible — and the average retune ratio is exactly the phasor's ratio. When it isn't, every crossfade splices a phase jump into the output, and the jumps do not average away: they bias the pitch. Period-locking the window is not a quality nicety; it is what makes the ratio true.

The fix: the window targets the smallest even multiple of the detected period that clears the minimum — 2 periods normally, 4 for high pitches whose 2 periods would be under 5 ms — so the taps always sit an integer number of periods apart. The clamp survives only as an outer bound. This is the cleanest example in the book of a defect no assert-on-internals test would ever catch: only an oracle — the detector listening to the output — could hear 5 cents.

Choosing the target

Scale mode: round the detected MIDI to a center note, then scan offsets −6…+6 for enabled pitch classes, keeping the candidate nearest to the fractional detected pitch (ties resolve to the smaller motion). A tritone of search radius suffices for any non-empty mask; an empty mask returns no target, and no target means a zero correction goal — the object never guesses. MIDI mode is simpler and blunter: the nearest currently-held note in absolute MIDI space, whatever the distance (clamped to ±12 semitones of actual correction). amount scales the goal before the glide — a fader on the distance itself.

The glide

The applied correction chases its goal through a one-pole with time constant speed (0 = assignment, the hard snap), evaluated per sample so the goal can move every hop while the glide stays silky. The ratio is then 2^(applied/12), computed per sample; the exp2 is cheap and the alternative — caching with edge cases — is not. The grain window rides its own 15 ms slew toward the period-locked target, and the detected period gets a third slew for the PSOLA backend's per-sample period input. Three small slews, no zippers, no special cases at the joins.

Three backends behind one seam

set_backend() swaps only the last stage; detector, mapper, and glide are shared state that survives the switch. Both alternate engines are constructed at prepare() (PSOLA sized to the deepest detectable period, the phase vocoder's FFT scaled to ~21 ms at any rate), so switching allocates nothing and is safe mid-audio; the incoming engine is cleared to silence first — a fade-in, not a splice of stale buffers. The ledger, at 48 kHz:

backendresynthesislatency
grainperiod-locked two-tap (in-kernel)≈ base delay + window/2, a few ms
psolatap::dsp::psola2 × 873 + 2 = 1748 samples ≈ 36 ms
pvoctap::dsp::pvoc (+ optional LPC formant trade)1024 samples ≈ 21 ms

The backend-parametrized scenarios feed all three the same 46-cent-sharp sawtooth and require the same landing (±6 cents at healthy level); a switching scenario hops between engines mid-signal and requires finite output and a correction that is still standing at the end. set_formant() forwards to the phase vocoder — PSOLA preserves formants by construction and the grain engine is waveform-preserving, so the flag deliberately touches one path.

Learning the key

The auto-key learner is thirteen doubles and a policy. Every voiced analysis adds 1 to its pitch class's histogram bin; every analysis multiplies all twelve bins by a leak chosen so the histogram forgets with a 60-second time constant. On demand — never on a schedule — the histogram is scored by Pearson correlation against the published Krumhansl–Kessler major and minor profiles at all twelve rotations, and the best of the 24 becomes the estimate, with the winning correlation as confidence. A mass guard withholds any estimate until roughly half a second of voiced material exists, so silence cannot have an opinion.

The design decision that matters is that the learner is advisory: autokey_estimate() reports and autokey_apply() adopts, but nothing in the audio path ever re-aims the targets on its own. This is a UI-safety argument, not modesty — a corrector that changes its own scale mid-phrase turns a wrong estimate into a wrong performance, and the person at the patch cannot undo what they never saw happen. Measured: a tonic-weighted D-major scale scores D major at 0.95 confidence; an A harmonic-minor melody scores A minor; reset withdraws the estimate.

The ledger

  • Allocation discipline. Everything sized at prepare(): detector frame and ring, both alternate backends, the grain buffer at the maximum window. After that the audio path allocates nothing; every setter either writes a double, flips a flag, or clears preallocated state.
  • Oracle-based testing. The scenarios measure the output's pitch with the independently-certified DspTap detector — the only kind of test that caught the period-lock bias — and use sawtooth, not sine, wherever PSOLA participates, per its documented material contract.
  • The maxtest. One assertion runs inside a real Max: unpitched DC in, exactly DC out — detector unvoiced, correction zero, complementary envelopes summing to one. It pins the whole "never guess" policy at unity gain in the shipping binary.
  • What the wrapper adds. Only plumbing: attribute forwarding, the atomic pitch handoff to a scheduler timer for the right outlet, and applykey writing back through the attributes so Max's saved state stays the source of truth.

Checkpoint

A per-sample corrector with a per-hop brain: worst-case geometry bought at prepare() so nothing ever allocates again, unpitched input relaxing to zero correction, and three slews smoothing every join. The war story is the period lock — two taps half a window apart are only honest when that half-window is an integer number of periods, and only an oracle test could hear the 5-cent lie. Targets are chosen, never invented; the glide is one pole and one exp2; the backends swap behind a seam that clears to silence; and the key learner watches, scores, remembers for a minute — and speaks only when spoken to.

The clipper in the loop: overdrive.h

The user-facing chapter claimed that tap.overdrive~'s gain tilts with frequency and that the tilt grows with drive — behavior a memoryless waveshaper cannot produce. This appendix derives the loop that produces it, shows why the obvious implementation of that loop is a stability bomb and how the file defuses it, and records the design decisions — shaper choice, asymmetry mechanics, oversampling versus ADAA — with the alternatives they beat.

The design brief was not a schematic (none is published for the Little Green Wonder, the listening reference): it was the class of TS-lineage feedback overdrives. The honest statement of the goal, from the project's handoff notes: the interesting part is not the transfer curve — it's the frequency-dependent gain and the softer, never-fully-flat knee that a feedback clipper gives you.

The topology, and what it must do

In a TS-lineage pedal the diodes sit in the feedback path of a non-inverting op-amp stage whose feedback network is frequency-dependent. Two consequences:

  1. The loop gain — and with it the effective clip threshold — varies with frequency: bass sees little gain and stays clean, mids see all of it.
  2. The output is input + limited feedback term: even at maximum drive the transfer's slope never reaches zero, because the clean input always passes.

overdrive.h models this with the minimal structure that keeps both traits:

w = shape( G·x − g_fb·LP(w) )      the clipper inside a lowpass feedback loop
y = x + w                          the unity clean path (non-inverting topology)

Signal-flow diagram: preamp and body pre-EQ into the oversampled region, where the gained signal meets a summing node, the shaper, and a lowpass feedback path; a unity clean path bypasses the shaper; DC block and body post-EQ follow at base rate

The whole kernel on one line. The red loop is the frequency-dependent gain; the amber path is why the transfer never flattens; everything inside the dashed region runs at the oversampled rate.

G is the drive gain (a dB sweep, +6 to +46). The lowpass LP (one-pole, corner 660 Hz) makes the fed-back signal predominantly low-frequency, so the negative feedback suppresses gain exactly where the pedal does. In the linear region (shape ≈ identity) the loop's small-signal gain is

w/x = G / (1 + g_fb·|LP(ω)|)

— at DC, G / (1 + g_fb); far above the corner, G. The file picks g_fb from a single voicing constant: g_fb = G/k_lf_gain − 1 with k_lf_gain = 2, which pins the low-frequency gain at +6 dB regardless of drive while the mids ride G all the way up. That one line is the measured headline — a bass-to-mid tilt of +5/+16.3/+17.2 dB at drive 0/0.5/0.9 — and it is the real-pedal behavior: turning up a TS makes the mids filthier while the low E barely moves.

Why the loop must be solved zero-delay

The naive discretization feeds back yesterday's lowpass state:

s    = G·x − g_fb·lp_state        // uses the previous sample's state
w    = shape(s)
lp_state += a·(w − lp_state)

That inserts a unit delay into a feedback loop — the same mistake as the Chamberlin SVF, with the same fuse. Linearize it: the state-to-state map has Jacobian J = (1 − a) − a·g_fb·shape′. At 48 kHz × 4 oversampling, a 660 Hz one-pole has a ≈ 0.021; at drive 0.9, g_fb ≈ 62. With shape′ = 1 (small signal — the quiet case!) J ≈ 0.979 − 1.34 = −0.36: stable, fine. But push g_fb higher — drive 1.0 gives G = 200, g_fb = 99 — and J ≈ 0.979 − 2.12 = −1.14. |J| > 1: the loop limit-cycles near Nyquist, audible as a parasitic whine that comes and goes with the signal level. A feedback clipper that oscillates when you turn it up is not a pedal, it's a bug report.

The fix is the house zero-delay move (svf.h's driven circuit, ladder.h's solver_fast): integrate the one-pole trapezoidally (TPT), solve the loop's linear part implicitly, then apply the nonlinearity and commit its output to the state. With the TPT one-pole v = (g·w + s)/(1 + g), substitute into the loop and solve for the node as if shape were identity:

w_lin = ( G·x − g_fb·s/(1+g) ) / ( 1 + g_fb·g/(1+g) )
w     = shape(w_lin + bias) − shape(bias)
v     = (g·w + s)/(1+g);   s ← 2v − s

No delay in the linear loop, so no delay-induced instability at any g_fb; and because shape′ ≤ 1 everywhere, the committed value only ever reduces the effective loop gain below the linear prediction — the approximation errs toward stability. At DC the solve gives w = G·x/(1 + g_fb) exactly, which is what makes the pinned-bass-gain arithmetic above exact rather than approximate. The kernel suite pins the consequence: after a full-drive, full-asymmetry signal stops, the output decays below 10⁻⁶ — no limit cycles.

The shaper: u/√(1+u²), and why not tanh

Three candidates from the brief, in the order they were rejected:

  • std::tanh — the reference softclip, and the expensive outlier: a transcendental call per (oversampled) sample that vectorizes badly.
  • Padé-style tanh approximations — cheap, but the usual forms are exact only on a bounded interval and go flat (or worse, retreat) beyond it — reintroducing the hard plateau this design exists to avoid, with a curvature discontinuity at the seam that aliases.
  • shape(u) = u/√(1+u²) — chosen: C∞ (no curvature seam to alias), strictly monotonic, asymptotic to ±1 but never flat, one multiply-add and one square root — which vectorizes as a reciprocal-sqrt instruction on every SIMD ISA this kernel targets, and reduces to a small LUT for a future fixed-point port.

The three candidate curves overlaid: the hard clip's corners, tanh's tighter knee, and the chosen rational curve's gentler, seam-free approach

The chosen curve reaches its asymptote more slowly than tanh — softer knee, lower-order harmonics — and unlike the hard clip it has no corner for the spectrum to pay for.

Asymmetry — the even-harmonic control the odd-only Jamoma curves structurally lacked — is a bias inside the shaper, output-corrected so silence stays silence: w = shape(u + b) − shape(b) with b = 0.5·asymmetry. At asymmetry 0 the whole path is an odd function and the measured H2 sits at the numerical floor (−151 dB); at 0.6 it is −26 dB and musically present. The correction term keeps the first-order DC out, but a biased clipper still rectifies: under signal it makes DC, and the feedback one-pole would happily integrate it. Hence the DC blocker after the clipper — y[n] = x[n] − x[n−1] + 0.9997·y[n−1], the Jamoma TTDCBlock constant kept for provenance — permanently in the path, not an option. (The original TTOverdrive instantiated that same blocker and then overwrote its output buffer without using it; the vestigial call was one of the tells, noted in the handoff brief, that the old code path was never going to be the base.)

Oversampling, not ADAA (for now)

Clipping generates harmonics without limit; everything past Nyquist folds back inharmonically. Two published remedies: oversample the nonlinearity, or antiderivative anti-aliasing (Parker et al., DAFx-16). ADAA is cheaper per dB of alias suppression, but its x[n] ≈ x[n−1] fallback branch is hostile to the branchless-SIMD constraint this kernel inherits from its embedded targets, and its difference quotient loses precision in single-precision float — a real concern for the fixed-point/f32 ports. So v1 oversamples: zero-stuff + 4th-order Butterworth anti-image up, matching anti-alias down — the ladder.h/svf.h resampler verbatim, self-contained per house rule. Factors 1/2/4/8, default 4×. Measured on a hard-driven 5 kHz tone: the folded seventh harmonic improves from −22 dB (1×) to −36 dB (4×) while the in-band harmonics stay within measurement error. ADAA remains the flagged experiment for after the voicing locks, so the comparison is apples to apples.

Output spectra of a 5001 Hz tone at drive 0.9, 1x overlaid on 4x: the 1x trace shows alias peaks standing tens of dB above the 4x floor

Every red peak standing above the blue mass is inharmonic fold-back the default 4× removes; the true harmonics (multiples of 5001 Hz) coincide in both traces. The dashed line marks the folded seventh harmonic the kernel suite pins.

The voicing layer, honestly labeled

Everything above is structure; the sound of the body control is a handful of constants (k_voice_* at the top of the file): the pre-clipper highpass corner sliding 40→320 Hz across the knob, the upper-mid bell at 1150 Hz (above the classic TS hump — the LGW's push sits higher), the +2.5 dB counterclockwise treble shelf, the fixed +1.5 dB mid seasoning. They produce the measured control shape (±10 dB at 100 Hz between extremes, +4 dB at the bell) and they are by-ear placeholders: the header says so, this book says so, and the numbers will move when the in-Max voicing pass against LGW demos happens. What will not move is where they live — all linear EQ outside the nonlinearity, because in the reference pedal that is what the Body knob is.

The parameter block is normalized on purpose

drive and asymmetry are 0..1, body is −1..+1; only preamp/output carry units (dB). The perceptual mapping (dB sweep of G, level compensation) lives inside the kernel, not in the knob range — so the parameters map directly to controllers, to live.dial, and to Q15/Q31 fixed-point registers on the Cortex-M targets this library's headers are written to reach. Parameters ride the standard per-sample linear ramps (default 20 ms); the derived coefficients — G, g_fb, the solve constants, the voicing biquads — refresh only on samples where a ramp actually moved, the same two-tier scheme as svf.h.

Everything in this chapter is executable: the loop math and stability claims are pinned by tests/overdrive_test.cpp (silence decay, tilt-grows-with- drive, even-harmonic emergence, DC blocking, alias improvement, determinism), every number is a cell in the verification notebook, and the measured figures are regenerated from the shipping kernel by book/figures/overdrive.py.

Wear as the stabilizer: tape_loop.h and discreet.h

Every regenerating loop in this library before these files made the same promise the same way: the loop is strictly contractive because feedback is capped below one (delay.h's k_fb_max = 0.99, the comb bank's calibrated ring time). tape_loop.h and discreet.h exist to make the opposite promise — regeneration at exactly 1.0, bounded anyway — and this appendix is the derivation of why that is allowed.

A shared header, by the house rule

The family needed the same four pieces twice (discreet.h and airport.h are both tape machines), and the reuse rule sorted them cleanly. Classes with state went into a shared header the way swing_vca.h was created for the drum family: tape::reel, tape::wow_flutter, tape::wear, and a tape::ramp that is a cited copy of delay.h's anti-zipper unit. Few-line expressions stayed copies-with-citation, as ever: the Hermite polynomial inside reel is the same read as delay.h, line for line, and says so; the saturator is not copied at all but included — vca::swing_shape, the shared swing-type stage, with the reason on the include line.

reel: one wrap, two topologies

A reel is position-addressed circular storage whose reads and writes wrap modulo a settable loop length, not the buffer size. That one decision lets the same class serve both kernels. discreet.h runs it as a delay line: loop length equals capacity, an integer write head advances forever (wrapped into range each sample — a bare long head would overflow LLP64's 32-bit long in half a day of audio), and the play head trails it by the loop span. airport.h runs it as a true loop: length set per piece, one free-running head, positions handed in raw because the reel does all modular arithmetic itself. A length change is deliberately a splice — content kept, positions re-wrapped — because that is what cutting tape does.

wow_flutter: periodic on purpose

The transport error is two sines — slow-deep wow, fast-shallow flutter — returning a read-position offset in samples, phases zeroed at prepare(). The periodic term is the dominant one in the tape-echo literature (Arnardóttir, Abel, Smith, AES 2008), but the deeper reason the stochastic term is a documented non-goal is testability: the wow promise is pinned by predicting peak pitch deviation in closed form (depth · 2π · rate, so 2 ms at 0.5 Hz ⇒ ±10.9 cents) and measuring it with the YIN oracle — 10.9 measured — and that oracle test only exists because two renders are bit-identical. Determinism was a design force here, not an afterthought.

wear: the boundedness argument

One pass of generation loss is three stages in fixed order: an exact one-pole darkening lowpass (1 − e^(−2πf_c/sr), the grm_comb.h map), the shared saturator swing_shape(v, d) = tanh(d·v)/d, and the normalized DC blocker. Each carries one clause of the proof:

  • tanh is bounded, so for any drive d > 0 the wear output can never exceed 1/d — whatever the loop has accumulated. That is BIBO stability at regen 1.0, unconditionally, from the saturator alone.
  • The DC blocker (pole 0.999, peak gain normalized to exactly 1 — the normalization grm_comb.h earned the hard way, chasing a +0.2 dB/s swell) kills the one frequency the lowpass would happily sustain forever with an offset attached.
  • The lowpass is strictly contractive above its corner and asymptotically transparent below it — which is not a leak in the proof but the musical contract: at drive 0 and regen 1.0 the sub-corner band sustains indefinitely, cleanly. The header calls this the Frippertronics contract and states it rather than hiding it.

So where delay.h proves stability by gain, this family proves it by shape: each pass survives because it is degraded. The pinned test drives regen 1.0 for ten seconds of ring and asserts non-growth — never decay, because decay would betray the contract just as surely as growth.

The doppler decision

discreet::machine gives loop_seconds an ordinary ramp and does nothing else, because nothing else is needed: moving a fractional read head is tape-speed doppler. A 0.5 → 0.75 s glide over half a second reads back an octave down mid-move (measured: 220 Hz, then re-lock within five cents) with no discontinuity, since position is continuous even where its slope is not. The rejected alternative — crossfading between two taps — would have hidden the machine, and hiding the machine is the one thing this kernel is for. The wow offset is clamped so the read can never cross the record head; at absurd depths on short loops the transport flattens against the clamp rather than wrapping, which the header files under honest limits.

A finding: the arithmetic agreed

The per-pass wear transfer is fully analytic — regen · |H_lp| · |H_dc| on the unit circle — so the notebook measured it the direct way: a two-tone burst (300 Hz under the corner, 6 kHz over it) recirculated at drive 0, each generation's tones read by Goertzel. Measured per-pass ratios: 0.292 and 0.890. Predicted: 0.292 and 0.890. Three decimals of agreement between a rendering kernel and a formula derived independently in the test is the cheapest kind of confidence this library knows how to buy, and both the test (with 15% and 5% tolerance bands it never needs) and the executed notebook carry the measurement.

The engineering ledger

The suite leans on four instruments. Analytic transfers wherever the path is linear (the per-pass darkening scenario asserts against the exact formula, both tones, both directions — highs die faster and lows barely fade, so the test cannot pass vacuously). Two-window RMS for long-run claims, inherited from the comb bank's swell story: regen 1.0 rings ten seconds and the late window may not exceed the early one. The YIN oracle for anything with a pitch: wow depth in cents against the closed form, the doppler glide and its re-lock. And bitwise assertions where the law is exact: mix endpoints, the first echo returning as literally the recorded impulse, two wow renders identical to the bit. The DC-step scenario checks the blocker's actual job — a held offset at regen 1.0 does not accumulate and the tail's mean returns below 0.02 — rather than a decay the contract never promised.

Checkpoint

One shared header, four blocks: a reel that wraps at the loop, a transport that is two deterministic sines, a wear stage whose tanh bound is the stability proof, and a cited copy of the house ramp. discreet.h composes them into the two-machine loop where regeneration legally reaches 1.0, loop moves are doppler because read heads are physical, and every claim is carried twice — discreet.ipynb executed, discreet_test.cpp pinned.

Free-running heads, one shared clock: airport.h

airport::loop_bank is structurally the smallest kernel in the family — a fixed array of loops, a stereo sum, no feedback anywhere — and that is what makes it interesting to read: nearly every promise it makes is structural, so nearly every test on it is bitwise. This appendix walks the file in code order and dwells on the one discipline that defines it.

loop_state: the multitap idiom with a reel in each seat

The bank is std::array<loop_state, k_max_loops> with an active count — delay.h's multitap shape, kept deliberately: per-index setters that silently no-op on a bad index, getters that return safe defaults, newly activated slots arriving at their stored settings. Each seat holds a tape::reel (its own worst-case buy — eight 30-second reels is ~92 MB of double tape, the family's largest allocation, stated in the header rather than discovered in production), a tape::wear used as a playback shade, a phase, a record flag, and three ramps (level, pan, darken).

The phase discipline

The load-bearing sentence in the header is "the phase is NEVER reset": recording starts wherever the head is, set_loops activates a loop with its head wherever it last was, a splice re-wraps the head modulo the new length without rewinding, and only prepare()/clear() — DSP restarts — may rewind. The reason is musical: in "2/1" the free-run is the piece, and any convenience reset (snap to zero on record, realign on length change) would quietly delete the composition. The pinned scenario earns the promise the blunt way: it fires a setter storm mid-render — level, darken, record, length, count — and then requires the click grid unmoved and the head advanced by exactly the samples processed. phase() exists as introspection precisely so that test could be written.

Record semantics

record is a gate, not an action: while on, the input replaces the tape at the integer head position, after the read — so you hear the previous generation under the head while punching, and one Hermite support point (two samples) of the old generation blends across the punch, which the header files under honest limits instead of papering over with a crossfade. No overdub-sum, because the provenance had none: each Airports phrase was recorded once. Freeze is the strong promise — record off, and two successive passes of the loop are required bit-identical. That promise is only possible because of the next decision.

The shade and its bypass

Per-loop darken reuses tape::wear with drive pinned at 0, as a static playback tone — deliberately not generation loss, because a frozen loop replays the same magnetic imprint every revolution and modeling wear on it would be dishonest physics. At the band ceiling (the default) the stage is bypassed entirely: not "flat enough", but not-in-the-signal-path, which is what upgrades the freeze test and the hard-pan test (a pan of −1 adds the loop's samples to the left bus unscaled) from tolerance checks to bitwise facts. Engaged, the shade is the exact one-pole from grm_comb.h, and the notebook measures a 6 kHz phrase through a 1 kHz shade at 0.169 of its transparent twin against 0.169 predicted.

composite_period_seconds

The lcm of the active loop lengths in samples, folded pairwise with a long long gcd, overflow detected before each multiply and reported as +inf. It is introspection, not DSP — but it is the piece's thesis as a number: 24000- and 30000-sample loops report exactly 2.5 s (and the pinned scenario also proves the rendered output repeats at 120000 samples and does not repeat at 60000), while seven airport-scale lengths overflow to infinity, which the header calls the point.

A finding: the raster before the assertion

The lcm scenario existed as an assertion first — bitwise equality of two 2.5-second windows — and it passed, which is exactly why it was worth plotting. The notebook's event raster (every return of loop A, loop B, and their sum on one timeline) made the same fact visible: the coincidence pattern audibly and graphically re-enters at 2.5 s and drifts everywhere short of it. The assertion pins the promise; the raster is what convinces a human the promise means something. The pair — one bitwise test, one executed figure — is this library's preferred way to hold a structural claim from both sides.

The engineering ledger

Almost everything here is exact, so the suite asserts exactly: bit-equality for freeze and for the lcm window, bitwise silence on the far bus for hard pans, phase() continuity to 1e−9 through the setter storm, and the splice law (0.9 of a 1 s loop re-wraps to 0.8 of a 0.5 s loop, never zero). The one measured tolerance in the file is the shade's analytic transfer at 20%, and the equal-power pan law needs no scenario of its own because the multitap chapter already pinned the center at 1/√2 to 1e−12 — same code shape, same law, cited rather than re-proven. Long-run behavior needs no stability test at all: there is no feedback path to go wrong, which is itself a fact the file's structure makes obvious enough not to test.

Checkpoint

A fixed bank of reels, one sacred free-running head each; record replaces and freeze is bitwise; splices re-wrap, never rewind; the shade bypasses to bit-transparency at the ceiling; and the composite period is the score's arithmetic made introspectable. The promises are structural, the tests are bitwise, and the executed raster in airport.ipynb is the human-readable proof that the structure composes.

Events, not audio: garden.h

garden::bed recirculates events where its siblings recirculate samples, which makes it the family's odd one out mechanically and its purest member conceptually: the wear-as-stabilizer inversion survives the abstraction jump intact, as arithmetic. This appendix walks the machinery — the ring, the split between planting and firing, the chime, the quantizer, the gardener — and the two contracts that had to be designed before they could be tested.

The event ring

Sixty-four fixed seats (std::array, nothing allocated at prepare() — this kernel buys no tape at all), each event a pitch, a velocity, a brightness, a position on the loop, and a plant-order sequence number. The sequence number exists for one policy: when the garden is full, the oldest live bloom yields to a new plant. The musical argument is stated in the header — a touch must always speak (rejecting input makes an instrument feel dead), and the oldest bloom has survived the most decay passes, so it is the quietest thing on the table; retiring it is the least audible edit available. The pinned scenario plants a distinctive high note, floods the ring with sixty-four more, and requires the first note's pitch measurably gone from the following pass.

Fire is not plant

note() does not sound a voice. It quantizes, seats the event at the loop's current position, and returns; the next process() sample finds the event's position under the playhead and fires it. The first draft did both — plant-and-fire in note() — and the loop fired it again one sample later, a double-trigger that fell out of the design the moment firing became the loop's exclusive job. One mechanism, two consequences: a plant sounds one sample late (inaudible, documented), and every sounding of every event goes through a single code path, which is what makes the return grid a testable promise. After each fire the event blooms: velocity times decay, brightness times soften, retire below floor — so a bloom lives exactly ceil(log(floor/velocity)/log(decay)) passes and the population converges no matter the planting rate. That is the stability theorem, and it is three lines of arithmetic instead of a saturator.

The chime

Four decaying mode doublets at the transverse-vibration ratios of the selected material — the free-free tube's 1 : 2.756 : 5.404 : 8.933 from the bars-and-tubular-chimes chapter of Fletcher & Rossing's The Physics of Musical Instruments (f_n grows as (2n+1)²), or the tuned bar's double-octave 1 : 4 : 10 : 20 from the mallet-percussion chapter — each mode a pair of sines split a fixed few cents, the doublet splitting of a real tube's degenerate mode pairs (same source), so the tail beats slowly instead of decaying like a lab sine. The ratio and haste tables are indexed [material][mode] and read at strike time, which is the whole implementation of the material switch: instant, allocation-free, and every live bloom re-voices at its next return. Each mode rides its own tr808::decay_env; decay times divide by ~ratio² (radiation damping grows with frequency), which makes the fourth mode a tens-of-milliseconds contact tick, and scale by √(440/f) per strike, so small high tubes ring shorter than long low ones. The upper modes scale with per-event brightness times strike hardness (a soft strike is a dull strike) and progressively steeply (b, b², b³), so soften strips the tick first and mode two by exactly its ratio. Mode levels sum to at most 1, so a chime is bounded by its velocity and the pool bound stays arithmetic; modes above 0.45·sr stay silent rather than aliasing.

The phase rule earned a refinement when the doublets arrived: a strike on a silent tube zeroes its phases — fresh initial conditions, so the pair starts aligned and its beat blooms identically at every return, which is what keeps per-return spectral measurements deterministic — while an audible steal keeps free-running phases and glides instead of clicking. The inharmonicity moved the pitch contract rather than breaking it: the upper modes clear quickly, so the YIN oracle reads each strike in its ring-down, and the scale-contract scenario still lands every off-scale plant on the scale within 20 cents.

The tube is the identity

Two more properties hang off each pitch, and neither touches the rng. A tube's upper modes sit up to ±3 cents off the ideal ratios — the fundamental stays true, because a maker tunes the fundamental — and the tube keeps a fixed seat on the stereo rack, spread scaling how far off center. Both are drawn by tube_unit, a stateless xorshift64* hash keyed by (fundamental-in-centihertz, index) — the metal_bank.h per-index idiom, index 0 the seat, 1..3 the mode scatter. Stateless is the load-bearing word: the gardener's seeded generator is never consumed, so the seed-triad contract survives intact, the rack is identical in every instance, and every return of a bloom rings from the same place with the same flaws. The pinned scenarios measure the scatter by scanning a Goertzel probe across the second mode (±0.25-cent steps resolve it), and the seat by left/right energy share: deterministic per pitch, different across pitches, bounded by the constants. The seat itself follows the phase rule — pan gains snap on a silent tube and slew ~10 ms through an audible steal — and the equal-power law is the √((1∓p)/2) form, so spread 0 makes the busses bitwise identical (also pinned).

Quantize at entry

The scale machinery is tune.h's 12-bit pitch-class mask idiom — the make_mask builder, the nearest-allowed search that never travels more than a tritone — copied with citation, not included, because tune.h reaches into tap::dsp for its detector and a garden should not link a pitch tracker to hold five scale presets. The masks themselves are plain public-domain scale theory, deliberately not any app's preset list. Quantizing at entry (rather than at fire) is the semantic choice: a scale change re-pitches nothing already planted, which keeps running gardens stable under live tinkering and makes the contract easy to state.

The gardener and the seed

The gardener is a wind model: once the idle threshold passes, strikes arrive on a calm/gust cycle driven by a small state machine — a gust catches 1 to 5 neighboring tubes (gust sizes it) with 30–280 ms between strikes, the clapper walking a few semitones per swing, and the following calm stretches with the gust just spent so the average rate stays near one strike per pass at any setting. Idle planting consumes the family RNG (tr808::white_noise, xorshift64*, the seed-folding and clear-reseeds contract) — and only idle planting does. That consumption discipline is load-bearing: the third leg of the seeded triad, "with the gardener disabled the seed cannot matter at all", is only true because a disabled gardener never touches the generator. The suite pins all three legs, plus the wind itself: at gust 1 some strikes tumble inside a gust, at gust 0 single strikes never come closer than the minimum calm (the scenario sets decay 0 so only the gardener's own strikes are counted — planted seeds recirculate, and returns are not wind). step_seq.h promises "no randomness anywhere"; this kernel is the deliberate counterpoint, and the triad is the bridge back to a reproducible test suite.

A finding: envelopes never reach zero

The return-grid scenario was first written the obvious way — the percussive test bell surely dies between returns, so the first nonzero sample after silence is the onset. It failed, instructively: decay_env's exponential tail crosses the 1e−12 hard-zero more than half a second after a "20 ms" decay, so there is no silence between returns, only −200 dB of not-quite. The fix was to stop pretending: an instant-attack bell, an amplitude threshold scaled to the expected return velocity, and a grid claim of "within 8 samples" — a sixth of a millisecond — with the comment explaining that a threshold on a sine sits a few samples into the cycle. The lesson is general for this library: exponential envelopes make "silence" a tolerance, and tests that assume literal zeros between notes are wrong even when they pass.

The engineering ledger

The suite measures the output, never the internals: fundamental ratios for the decay staircase (0.5 ± 0.05 across four returns — the fundamental, because hardness makes whole-strike peaks fade faster than velocity, then active_events() == 0 and the render below 1e−6), a strictly-decreasing Goertzel sideband for softening, YIN for the scale contract, the seeded triad rendered three times over, and structural bounds exercised at their edges — sixty-five plants against sixty-four seats, thirty-two notes against sixteen bells, finiteness and the k_voices amplitude bound under sustained stealing. The two introspection counts (active_events, active_voices) exist, as phase() does next door, so those scenarios could be written against public surface.

Checkpoint

A fixed ring of events fired by a loop counter into a fixed pool of modal wind chimes: plant and fire kept strictly apart, wear as per-pass arithmetic (decay, soften, floor) with convergence as its theorem, scale masks copied from tune.h and applied at entry, tube identity (material voicing, mode scatter, stereo seat) as stateless hashes so nothing generative leaks into the audio path, and a gardener whose RNG discipline makes generative behavior compatible with a bit-exact test suite. Third costume, same inversion: the system stays bounded because everything in it is always fading. Every claim lives twice — garden.ipynb executed, garden_test.cpp pinned.

Composition, not construction: tapecho.h

This is the shortest appendix in the book, and that is the point of it.

tape_loop.h was written for the Eno family — one shared header holding a reel, a transport, and a wear stage, factored out because discreet.h and airport.h needed the same four pieces twice. The claim implicit in factoring it that way was that it is a library: machinery that a machine nobody had written yet could be built out of. tapecho.h is the test of that claim, and the result is worth recording precisely, because "we extracted a shared header" is easy to say and rarely checked.

The result: tape_loop.h needed no changes at all. Not a new method, not a widened clamp, not a friend declaration. A tape echo — a different topology, a different number of read points, a different stability regime — composed out of it exactly as shipped.

What the file actually contains

Two classes and no DSP that was not already in the library.

head is a read position with three ramps: a ratio along the tape path, a level, and a pan. Its read() takes a reel it does not own, the motor span, and the shared transport offset, and accumulates a panned contribution onto the stereo busses. It is a component in the airport.h sense — a piece the monolith is made of, reachable for testing — but honestly labeled as not standalone-external material: a head without a reel is not a machine, it is an index. That distinction is worth keeping straight, because the components chapter's lesson ("the monoliths were monoliths by accident") can be over-applied. Some seams are real and some are arithmetic.

machine owns one reel in delay-line topology, one tape::wow_flutter, one tape::wear, and four heads. Its process() reads the heads, applies the regeneration cap, writes the record head, and mixes. There is nothing else in it.

The geometry: one motor

Each head's delay is span_samples * ratio - offset, where offset is the transport error. Two decisions hide in that one line.

The first is that span is defined as the delay of a ratio-1.0 head rather than as "the delay time", which is what makes the motor a motor: one multiply per head and the whole layout scales together, as a tape speed does. The alternative — per-head absolute times — would have made a speed change into four coordinated parameter moves and lost the doppler for free.

The second is that offset is subtracted once, shared by every head. That is physically right for a single transport (one capstan error displaces the whole tape path) and it is also the cheap answer, so it is worth saying plainly that the per-head phase differences of a real multi-head transport are not modeled. It is a documented limit, not an accident.

The stability inversion, one step further

machine/tape.md derived why discreet.h may run regeneration at exactly 1.0: wear's saturator is bounded by 1/drive, so the loop is bounded no matter the gain. That derivation does not stop at 1.0 — nothing in it does. So this kernel lets regeneration reach k_regen_max_driven (1.5), and the tape is bounded by |in|max + regen/drive at any setting.

The subtlety is the boundary. That guarantee exists only while the saturator is engaged, and drive is a ramped parameter a performer can take to zero mid-howl. At drive 0 the wear path is exactly linear with |H| ≤ 1, so regeneration above 1.0 would grow without bound. The kernel therefore computes the cap per sample from the current drive:

const double regen_eff = std::min(regen, (drive > 0.0) ? k_regen_max_driven
                                                       : k_regen_max_linear);

Not in the setter — in the audio path, because drive moves during performance and a setter-time decision would be stale the moment it mattered. The stored target keeps its high value, so pulling drive to zero lands the loop at 1.0 and restoring drive brings the howl back. That asymmetry between target and effective is the one piece of state in this kernel that is not obvious from the header's public surface, which is why it is written down twice: here, and in the file's own banner.

Why the null test is the important one

The suite's load-bearing scenario neutralizes the tape — no transport error, no regeneration — and asserts that a one-head echo is bitwise delay.h's Hermite multitap.

It is bitwise rather than approximate because nothing was reimplemented: both paths compute time_ms * 0.001 * sr the same way, both clamp at the same 2.5-sample Hermite floor, both evaluate the same polynomial at the same fractional position, and both apply (pan + 1) * 0.25 * π to the same k_pi. The multiply by a ratio of exactly 1.0 and the subtraction of an offset of exactly 0.0 are both exact in IEEE-754, so the arithmetic does not merely agree — it is the same arithmetic.

That is what makes the test meaningful. An approximate null test would pass just as happily over a second implementation that happened to be close. A bitwise one only passes if the shared code is genuinely shared, which is the proposition on trial. The notebook runs the same comparison across the C ABI so the claim also holds at the boundary the externals cross.

One consequence worth knowing when reading the test: at pan 0 the two busses are not bit-identical to each other, because cos(π/4) and sin(π/4) differ by one ulp in IEEE-754 doubles. That is inherited from delay.h's pan law, it is the same in both objects, and it is exactly why the null test compares each bus against its counterpart rather than comparing left to right.

Checkpoint

Two classes, no new DSP, and a shared header that did not move. The motor geometry buys varispeed with one multiply per head; the regeneration cap lives in the audio path because the thing it depends on is performed; and the null test is bitwise because being bitwise is the only version of that test that proves anything.

Dice you can replay: stammer.h

Most kernels in this library are hard to get wrong quietly: a filter with a bad coefficient sounds bad. A stutter is not like that. It has three interacting integer clocks — a grid countdown, a slice origin, and a playback head — and if any one of them is off by a sample the object still sounds fine. It stutters. It grooves. It is just wrong in a way no amount of listening will surface.

So the interesting content of this appendix is not the DSP, which is a buffer and some dice. It is how you pin three clocks at once.

The pinned-dice identity

Every random draw in the kernel has a setting at which its outcome is forced. Fire probability 1 always fires. divisions 1 always picks the whole step. repeats 1 always plays one pass. reverse 0 never reverses. jump 0 never reaches back. fade 0 leaves the material alone.

Set all six and the machine becomes deterministic regardless of the seed — and what it must then be is not a vague "sensible output" but a specific, checkable thing: exactly a one-step delay. At each grid point it grabs precisely the step that just went past and plays it once, so

y[i] == x[i - step + 1]

for every sample after the first grid point, bitwise.

That single assertion is worth more than three separate off-by-one tests, because it fails if the grid countdown fires a sample early, if the origin arithmetic reaches one sample too far back, if the playback head starts at the wrong index, or if any two of those are wrong in ways that would cancel in a looser test. It is also cheap to reason about, which matters: a test you cannot re-derive on a whiteboard is a test you will eventually delete instead of fixing.

The + 1 in that expression is not a fudge. The write head holds the next write position, so after recording sample i the newest available sample sits at position i, and a slice of length step grabbed at grid point k·step reads positions k·step + 1 - step upward. Getting that constant right by derivation rather than by nudging until the test passed is the whole discipline; a test you tune to the implementation pins nothing.

Two smaller identities sit alongside it. With reverse 1 the same grab reads end-first, so y[k·step + j] == x[k·step - j] — the mirror of the first, which catches a reversed-index off-by-one that the forward test cannot see. And with repeats above 1, every output sample must be either a fresh grab or a bit-exact copy of the block one slice-length earlier; nothing else is legal, because a slice in flight is never interrupted. That invariant covers the repeat machinery without needing to know how many passes the dice chose.

The draw order is part of the ABI

maybe_fire() draws in a fixed order: fire, division, repeat count, reach-back, then the first reverse coin. Each subsequent repeat draws its own reverse coin as it starts.

That order is not an implementation detail — it is what "a seed is a performance" means. Reordering two draws, or adding a draw in the middle, silently changes every render anyone has ever made with a given seed. garden.h established the same discipline for the gardener; this file inherits it, and the comment above the function says so in as many words so the next person to add a parameter knows to append rather than insert.

The disabled case is the sharp end of it. At density 0 the function returns before drawing anything:

if (m_density <= 0.0) {
    return; // the dice are never rolled, so the seed provably cannot matter
}

The lazier version — draw, then compare against 0 and fail — behaves identically to the ear, and would be indistinguishable in almost any test. It would also consume one number per grid point, so the seed would matter: switch density off and on again, and the stream is somewhere else. The garden made this a family contract, and it is pinned here by a test that runs two different seeds at density 0 and requires bit-identical output.

Reading from the ring, and what it costs

A slice does not copy its material. It stores an origin and reads from the capture ring as it plays.

The alternative — memcpy the slice into a private buffer at fire time — would be more obviously correct, and it is what a first draft wants to do. It was rejected because it is a burst copy in the audio thread: half a second of slice is 24,000 doubles moved inside one process() call, a spike that does nothing for 47,999 other samples. Reading from the ring costs nothing extra.

The price is a real failure mode, so the header states it: if a repeat train outlives the buffered history — repeats · length + jump beyond max_history_ms — its tail reads fresher material as the write head laps the origin. The kernel clamps the slice length against the bought capacity so it can never read outside the buffer, but it does not and cannot prevent a long train from being overtaken. Sizing the history is the caller's job, and the object argument exists for exactly that.

The envelope, and why the dip stays

Flanks are raised sine, computed to be exactly 0 at both edges and exactly 1 across the plateau, clamped per slice to half the slice so the two flanks never overlap. Repeats are sequential, not overlapped, so each junction dips to zero rather than crossfading.

Leaving it that way was a decision. An equal-power crossfade between consecutive passes is easy from here — the pieces are already in the family — and it would smooth exactly the articulation that makes a stutter read as rhythm. The dip is the transient the ear locks onto. So the file documents it as intentional, and the test asserts the exact edges, so nobody later "fixes" the dip and quietly turns the object into a tremolo.

Reuse, and the component that is not a component

The capture is a tape::reel in delay-line topology — the same class the tape echo uses, the same class airport.h runs as a true loop. Reads here are always at integer positions, so the family's Hermite read reduces to an exact sample fetch; that is slightly more arithmetic than an integer index would need, and it is kept anyway because it is one code path and because a rate-varying sibling (tap.scrub~, planned) needs precisely this.

The randomness is tr808::white_noise, the family's seeded xorshift64*, reached through swing_vca.h. Nothing new was written for it.

Which leaves the split: capture and slicer are separate classes under a thin machine, per the family's components-first habit. But a slicer needs a capture to mean anything, so — like tapecho.h's head, and unlike airport.h's loop — it is documented as a component for composition and testing rather than a candidate for its own external. The components chapter's lesson is that seams often already exist and only the monolith can reach them. The corollary, which is easier to forget, is that not every class boundary is a seam.

Checkpoint

Three clocks, pinned by one identity: force every die and the machine must be exactly a one-step delay, bitwise. A fixed draw order, because that is what makes a seed a contract, and an early return at density 0 so a disabled generator provably cannot consume its stream. Ring reads instead of a burst copy, with the failure mode written down rather than papered over. And an envelope dip that is deliberate, tested, and therefore safe from being helpfully removed.

Two stages and a knee: fuzz.h

This appendix is mostly about two mistakes, because the DSP itself is a published recipe followed closely and there is little to explain about it that Yeh, Abel and Smith's DAFx-07 paper does not explain better. What is worth recording is what went wrong on the way, since both failures are the kind that recur.

The recipe, briefly

Two stage objects, each a conditioning highpass, a gain, a memoryless curve, and an equalization lowpass — the paper's cascade, twice. A tone section of three RBJ biquads outside the nonlinearity. A pedal that runs the pair inside an oversampled region, DC-blocks, voices, and trims.

The curve is shape(x, k) = tanh(kx)/tanh(k), chosen from the family the paper itself compares against a tabulated diode DC curve (tanh, arctan, a tanh approximation). The normalization matters more than the choice: dividing by tanh(k) fixes the output at full scale for unit input at every knee, so edge changes the shape of the corner without moving the ceiling.

Mistake one: small-signal gain compounds

The tanh family's slope at the origin is k/tanh(k). It is greater than one and it grows with the knee — 1.7 at knee 1.6, 2.1 at knee 2, 12 at knee 12.

In one stage that is a curiosity you can absorb into the gain mapping. In a cascade it multiplies. The first implementation put a fixed ×2.2 in front of a knee-3 curve, so the second stage's effective small-signal gain was about 6.6, and with the drive floor at +6 dB the limiter was already saturated with the gain knob at zero. The measured harmonic-to-fundamental ratio was 0.401 at gain 0 and 0.408 at gain 1: the knob did essentially nothing.

The reason this is worth a paragraph is that it is inaudible as a bug. The object sounded like a distortion pedal at every setting, because it was one at every setting. Only a swept measurement showed the knob was inert. The fix was to lower the drive floor below unity (−12 dB) and the second stage's fixed gain to 0.5; the ratio now runs 0.010 → 0.358.

The general lesson, stated for the next cascade someone builds here: a waveshaper's small-signal slope is part of the gain structure, and if the curve family's slope depends on a user-facing parameter, that dependence propagates to every stage downstream of it.

Mistake two: the house oversampler, and a hypothesis that was right

The oversampling chain in tap.ladder~, tap.svf~ and overdrive.h is zero-stuff plus a 4th-order Butterworth, cut at 0.45 of the base rate normalized to the oversampled rate. This file started as a copy of it.

Measured, that was wrong here: alias energy at the fold frequencies came out worse at 4× than at 2× (1.7e-2 against 2.8e-3). Twenty-four dB per octave leaves content just above the base Nyquist barely attenuated, and a higher factor pushes more clipper-generated content into exactly that band before decimation. Moving to 8th order improved 4× about sixfold.

It did not fix the ordering. An earlier draft of this file recorded that as an open question with one hypothesis ruled out and one surviving:

  • Ruled out: numerics. At 8× the filters are cut at 0.056 normalized, where biquad poles crowd the unit circle. Tested by running the cascade's impulse response out to 400,000 samples at each factor; it decays cleanly to denormal every time.
  • Surviving: imaging. Zero-stuffing by N leaves N−1 images for a single filter to suppress; residual images entering a nonlinearity intermodulate with the signal into products that are not harmonics of the input, which is precisely what the probe measures, and there are more of them at higher N.

ondes.h then supplied evidence for the survivor without being built to: same 8th-order chain, comparably hard nonlinearity, but a source with nothing zero-stuffed on the way up, and its sequence never reversed.

Acting on it settled it. The chain is now one 2× stage per doubling, each filtering at 0.225 of its own operating rate — a corner that never tightens however deep the cascade goes, which is the whole difference. Same probe, same material, only the resampler changed:

toneold 4×old 8×new 4×new 8×
3733 Hz7.4e-41.8e-32.1e-52.2e-5
4409 Hz7.9e-42.1e-33.7e-73.7e-7
5171 Hz9.2e-42.3e-32.0e-71.9e-7
6421 Hz5.8e-41.3e-33.9e-73.6e-7
9337 Hz2.9e-46.2e-51.8e-62.4e-8

The worst step-up past 2× is a ratio of 1.017 — flat, where the old chain ran up to 3× worse per doubling. The 2× column is unchanged in both, as it must be: one doubling is one stage either way, and that it is unchanged is the best available check that nothing else moved.

The cost is 3.16 % of a core at 8× against 3.02 % before. The filters are cheap next to the clipper they surround, which is worth knowing in advance next time this trade looks expensive.

Mistake three: one tone is not a sweep

Every number in the two sections above — the 4th-order finding, the reversal, the "2× is best" default that shipped — came from a single test tone at 3733 Hz. The tone was chosen carefully, for good reasons that are still good: it does not divide the sample rate, and its folds land where nothing else lives. It was still one tone.

Swept properly, 2× does not merely fail to be best. It collapses above about 6 kHz, and at 10499 Hz it measures worse than no oversampling at all (1.7e-1 against 1.5e-1), because the clipper's low harmonics already exceed the base Nyquist there. The default that shipped was safe only for material that stays below 6 kHz.

This is the same error as the two in the next section, one level up: those are about choosing a bad probe, this is about choosing too few. A probe that is correct at one point on the input domain tells you about that point. The fix is not cleverness, it is a for loop over tones, and it costs seconds.

Two properties of this particular probe bound where the loop can go, and both are now written down next to it: it only measures folding at all above about 3 kHz, since below that harmonics 8–13 are still under Nyquist and it reads real harmonics instead; and tones that are simple rational multiples of the sample rate stack folds on top of each other or put one exactly at Nyquist, where it reads nonsense.

Two ways to measure aliasing wrong

Both were committed before being caught, and both are recorded in fuzz_test.cpp because they are easy to repeat.

Choosing a tone that divides the sample rate. The first alias test used 3 kHz at 48 kHz. Every harmonic of 3 kHz folds back onto another harmonic of 3 kHz, so every alias hides exactly underneath legitimate content and the probes read nothing at all. The test passed happily while measuring noise. 3733 Hz puts the folds where nothing else lives.

Probing too close to the fundamental. Two of the original probe frequencies sat a few hundred Hz from a full-scale tone. What they measured was the window's spectral leakage — around 1e-3, which swamped the aliasing underneath it. Probes have to be far enough out that leakage from the loudest component is below the thing being measured.

Checkpoint

A published cascade, followed closely. One curve whose normalization keeps the knee from becoming a volume control. A gain floor set below unity because small-signal slope compounds across stages — the bug that sounded fine. An 8th-order oversampling filter because the house 4th-order one measured worse, and a cascade of 2× stages because a single zero-stuff by N was what made bigger measure worse — the imaging hypothesis, recorded as open here for two waves, then confirmed by acting on it. And a default of 4× rather than 2×, because the 2× default had been generalized from one test tone and collapses above 6 kHz.

One tape, two read patterns: scrub.h

When stammer.h shipped, its header made a promise in its own limits section: slices play at ±1 rate, a performable pitch-bending playhead over live capture is a different object, and sharing this capture is the plan. This file is that promise being kept, and it is worth recording that the sharing turned out to be literal. scrub.h includes stammer.h and uses stammer::capture itself — one tape_loop.h reel under an advancing write head — rather than keeping a second copy of the same idea.

The only thing the stutter had to grow was capture::read_frac, a fractional Hermite read. Its ±1-rate slices never needed one.

The rest of the file is two classes: head, the grain scheduler, which owns the grain pool, the hop clock and the spray dice and reads a capture it does not own; and machine, which is one capture, one head, the freeze gate, the drift and the balance. Same parts-then-composition habit as tapecho.h and stammer.h, for the same reason: head is a read pattern, not a machine, so it is testable and composable without being an external.

The defect that measurement caught and nothing else could

The first cut anchored every grain at the position. That is the obvious thing to do — the position is where the user is pointing — and it is wrong in a way that is genuinely hard to hear.

Here is the mechanism. If every grain's origin is write_head − lag, then origins advance at the write head's speed, which is exactly 1. Each grain then plays from its origin at rate. Inside a grain the pitch is correct. Across grains the average read rate comes back to 1, because the origins reset it every hop.

So a steady tone comes out at its original pitch, with a comb of grain-rate sidebands around it. The pitch knob did not transpose. It added texture, and texture is what you expect from a granulator, which is exactly why no amount of listening was going to find this.

The fix is a phase-continuous read head: the origin advances at rate, and is pulled back toward the position only once it has wandered more than ±1.5 grains. Every pull-back is a splice, which is the cost, and the bound is chosen by sweep rather than taste. Band energy retained around the transposed pitch, at wanders of ±0.5 / ±1 / ±2 / ±3 / ±4 grains:

wander (grains)meanworst
±0.50.9330.716
±10.9580.820
±20.9650.874
±30.9900.918
±40.9930.940

Flat past 3, and every extra grain of wander is a grain of position error, so k_wander_grains = 3.0.

The unity case is special-cased to zero error rather than accumulated, which is what keeps the null exact: at rate == 1 there is nothing to wander from.

Measure the band, not the bin

This is the second thing worth carrying out of this file, and it nearly inverted the conclusion above.

A single-bin probe reads the fixed kernel as badly broken. The splices spread the transposed partial into a comb a few hertz wide; a rectangular-window Goertzel sitting on one line saw 0.02 where the band figure was 0.43. Had that been the first measurement taken, the fix would have looked like the bug.

Measured properly — energy in a ±15 Hz band around the transposed pitch, against the same band of a perfect shifter — 98.8 % lands where it should, worst case 91.7 %. What the splices cost is concentration, not pitch: 92.0 % as focused as a clean shift, 75.0 % at worst.

The general rule, stated for the next time someone here measures a pitch-shifter: if the process can smear a partial, a single-bin probe is measuring the smear, not the partial. Integrate a band wide enough to contain the artifact you already know about.

And then, immediately, the same mistake in its other half. The comparison against tap.pitchaccum~ used that ±15 Hz band unchanged across the whole sweep — but ±15 Hz is about 115 cents wide at 220 Hz and only 26 cents at 932 Hz, so at the top of the sweep the probe was again narrower than the process it was measuring, and it produced two readings of 0.0001 and 0.0006 that were recorded as near-total cancellations of a shipped object. Widened to a constant 3 %, they read 0.63 and 0.85 and no cancellation exists. The retraction and what survives it are issue #33.

So the rule has a second half: a band wide enough in the units the process works in. A pitch shifter works in cents. A fixed hertz window is a different width at every pitch, and the place it is narrowest is exactly where a shifter's error is largest.

Two related mistakes are recorded here because both were committed:

  • Analysing mostly silence. The first wander sweep ran 1 second of material with a 900 ms position lag, so most of the analysed window was tape that had not been written yet. Extended to 3 seconds, analysing the last third.
  • Feeding a discontinuity into the test. A slew test drove the object with a sine and then, mid-test, called process(0.5) with a literal DC sample to change a parameter. That step was an input transient, and the 0.48 jump it produced was the test's own fault. Continuous tone index, and the same bug was then fixed pre-emptively in diffuseur_test.cpp.

The null, and the arithmetic that makes it exact

Hann satisfies constant-overlap-add at hop = size/overlap, so the window sum is exactly 1 at overlap 2 and above, and normalization is 2/overlap so the level holds across settings. With pitch at unity, spray at zero and the position on a whole sample, the object is the input delayed to 4.4e-16.

It is exact only when size divides evenly by overlap, because the hop is an integer number of samples; otherwise a small periodic ripple survives in the window sum. It is inaudible at musical sizes, and it is why the null test chooses the numbers it does (480 samples of lag, 96 of size) rather than round milliseconds.

The mix control needed the same care as the diffuseurs' — an equal-power blend written as cos/sin does not return exactly zero at the endpoint, and a wiring null that reads 6.1e-17 instead of 0 is not a null. Both ends are short-circuited exactly.

The grain pool starves rather than steals

Shrinking size sharply while grains are in flight can leave every slot busy at the moment the next grain is due. That grain is dropped, not allocated by stealing a slot from a grain mid-window, because a steal cuts a Hann window in half and clicks. The audible cost is a momentary dip, bounded by the pool being two slots deeper than the maximum overlap.

A limit that is not fixed, on purpose

A grain born lag samples behind the write head and playing at rate r reaches lag − size·(r−1) behind it by its end. Transpose up with the position near the live edge and the grain's tail runs off the front of the tape into the oldest material.

Nothing clamps this. Clamping would silently bend the pitch to keep the grain in bounds, which is a worse failure than the seam — the object would stop playing the interval you asked for and never say so. The constraint is documented (keep the position at least size·(rate−1) back) and left to the player.

Checkpoint

One capture, shared literally with the stutter, plus one fractional read that the stutter did not need. A phase-continuous read head, because anchoring grains at the position quietly cancels the transposition — the defect of this file, invisible to listening and obvious to a sweep. A wander bound measured rather than chosen. And a measurement lesson worth more than the kernel: a single-bin probe on a smeared partial reads the fix as the bug.

Driven, not struck: diffuseur.h

This file exists because a plan was wrong in a useful way. The Ondes family plan said the diffuseurs would inherit garden.h's modal machinery, and they do — mode ratios, doublet splitting, per-mode decay. What it did not say, and what reshaped the file, is that a diffuseur is driven. There is no trigger here and no decay_env. The input excites the body continuously and the body rings at its own rates, which is grm_comb.h's situation rather than the chime's.

Five classes: mode, plate, sympathetic, harp, transducer, and two cabinets over a shared cabinet base. Nothing else.

Unit peak gain, and everything it saves

mode is the constant-peak-gain two-pole resonator (Steiglitz; Smith, Introduction to Digital Filters): poles at radius R, zeros at ±1, and b0 = (1 − R²)/2.

That choice pays three times, and it is worth spelling out because it is the reason this file has almost no defensive code in it.

  • Peak gain is 1 at any Q. So a bank of weighted modes is bounded by the sum of its weights. The plate's eight weights sum to exactly 1, which means the body cannot output more than its input, and there is no limiter anywhere in the file.
  • Changing decay does not change the level. With a plain two-pole resonator, moving R moves the peak gain, so a decay knob is also a volume knob. Here it is not.
  • The zeros at ±1 are exact nulls at DC and Nyquist. So there is no DC blocker on the body either. It cannot accumulate one.

None of that is novel — it is a textbook resonator used for the reason the textbook gives — but the cumulative effect on a file that runs sixteen of them plus twelve delay loops is large.

The order is the argument, and it is a bitwise test

The instrument's signal reaches the transducer first, and the transducer's motion excites the body. So the nonlinearity is upstream of the resonator.

That is the central design claim of the file, so it is pinned rather than described: a scenario builds transducer → plate by hand and checks that a whole metallique is bitwise identical to it, and that the reverse wiring — resonate, then distort — differs by 28 % of peak.

Getting that null to be actually bitwise took one fix. cabinet::blend is an equal-power crossfade written with cos/sin, and cos(π/2) in double precision is 6.1e-17, not 0. A wiring test that reads 6.1e-17 has not demonstrated identity; it has demonstrated approximate identity, which is the thing the test exists to distinguish from. Both ends of the blend are now exact short-circuits: mix 0 returns the dry input bit for bit, mix 100 returns the wet.

The transducer's bound is 2/saturation, not 1/saturation

The moving-iron model squares the drive, and vca::swing_shape bounds the result at 1/saturation. The obvious test — output stays under 1/saturation — failed at 1.49 against a bound of 1.25.

The test was wrong, not the code. A hard-driven squared law produces a nearly-constant positive waveform: it sits up near the ceiling and dips toward zero. Removing its DC recentres that, so the excursion below the mean adds to the excursion above it, and the worst-case swing after the DC blocker is up to twice the saturator's own bound.

The corrected bound is 2/saturation, documented in the header, and the scenario now asserts both sides of it — greater than 1/sat, less than 2/sat — so the test still catches the saturator disappearing entirely.

The general shape of this mistake is common enough to name: a DC blocker after an asymmetric nonlinearity is not free. It does not just remove an offset; it converts an offset into headroom you have to have.

A measurement that measured its own edges

The palme's selectivity scenario drives the board with a tone and measures what is still ringing after the tone stops. First version: switch the tone on, switch it off, measure the tail. It failed — 3.66× selectivity against the 4× asserted — and the failure was real but not about the strings.

Switching a tone on and off is a step, and a step is broadband. It excites every string on the board, so the tail contained twelve strings ringing regardless of what frequency had been played. The measurement was reading its own edges.

Fading the drive in and out over 250 ms removes the step. Every one of the twelve strings then passes, with the worst at 4.4×. The same fade is what the book figure uses, and the figure's caption says so, because a reader reproducing it without the fade will get the wrong answer.

Twelve strings

Widely copied hobbyist build pages describe the palme as two banks of twelve strings. The peer-reviewed source (Wijnand, Boutin, Jossic & Maniguet, Forum Acusticum 2023) says twelve, and k_strings = 12 with a comment saying which source won and why.

Their tuning is not published anywhere found, so it is a parameter rather than a constant, and the header says that too. Guessing a tuning and hard-coding it would have been the same category of error as the twenty-four.

Where recreation begins

The instruments, their dates, their excitation and their transducer type are peer-reviewed. The modal data is not — no ondes-specific measurement of either body exists in any of the four sources read — so the plate uses Fletcher & Rossing's free circular plate (Rayleigh's Chladni ratios at Poisson 0.3) and the strings use the harmonic series.

This is stated in the header at the top rather than in a limits section at the bottom, because it changes what the object is: a recreation of the general physics, not a model of Martenot's instruments. The same applies to asymmetry and saturation — the source establishes that the moving-iron driver is nonlinear and that Thiele–Small does not describe it, and then does not hand over a curve. Those two coefficients are voiced by ear and labelled as voiced by ear.

A diffuseur with both at 0 is a linear resonator and is missing a real stage. That is a choice the caller may make, and the header says so rather than forcing a minimum.

Checkpoint

A textbook resonator chosen for three properties that between them remove the limiter, the DC blocker and the decay/level coupling. A bitwise null that pins the transducer upstream of the body, which required making an equal-power blend exact at its endpoints. A saturator bound corrected from 1/sat to 2/sat because a DC blocker after an asymmetric nonlinearity buys headroom, not just centring. A selectivity test that had to stop measuring its own on/off step. And a provenance line drawn where the published sources actually stop.

A citation, an identity, and a sign: ondes.h

Three classes — triode, detector, voice — and three things worth recording about how they got here. One stage turned out to need no design decisions at all. One approximation turned out to be an exact identity. And one sign error made a distortion knob run backwards.

The circuit is Najnudel, Hélie, Roze & Boutin, "Simulation of an ondes Martenot circuit", IEEE/ACM TASLP 28, 2651–2660, 2020, modelling instrument No. 169 as five port-Hamiltonian stages. This file is not that: their full solve runs at 768 kHz and their plugin costs 85 % of a laptop core. What it takes from them is their own published reductions plus their published component values, and the header says which is which.

The tube is a citation, not a design

The plan framed the valve stage as a choice: a published grid-conduction curve, or the tanh family with an asymmetry bias voiced by ear. It is neither, and finding that out took nothing more than reading the paper properly.

The paper names a tube model — the enhanced Norman Koren model (Koren, Glass Audio 8(5), 1996, with Cohen & Hélie's grid-current branch, AES 129, 2010) — writes out its three equations, and publishes parameter sets in Table II fitted to the actual valves in ondes No. 169, together with each stage's supply voltage, cathode resistor and plate load.

So there was nothing to voice. k_6f5, k_6c5, k_2a3, k_op_demod, k_op_preamp and k_op_power are Table II transcribed, and the header says they are the citation.

A stage is then the static solution of ipc(vpc, vgc) = (Vbias − Vk − vpc)/Rp on the load line, with cathode bias Vk = Rk·Ipc found at the quiescent point. That is a memoryless nonlinearity in exactly the DAFx-07 sense, which matters for a practical reason: tabulating it is not an approximation of the model, it is the model. The table is rebuilt on a tube or operating-point change and read with linear interpolation, so the audio path costs a lookup rather than a root find.

The published points bias sanely — the 6C5 demodulator lands at Vk 2.70 V, Vp 86.5 V, Ip 2.70 mA, gain 4.86 — which is its own small confirmation that the transcription is right.

The sign that made the drive knob run backwards

The stage must invert, as a real common-cathode stage does, and this is load-bearing rather than cosmetic. The valve's asymmetry acts on whichever side of the waveform reaches its grid. An early cut normalized the output by the signed small-signal gain, which quietly un-inverted the stage, so the curve's lopsidedness landed on the wrong half of the waveform.

The symptom was unambiguous once measured: turning drive up reduced total harmonic content. A distortion control that gets cleaner as you push it is not a subtle bug, but it is only visible in a sweep — at any single setting the object sounded like a valve.

Two changes fixed it. The curve is now the true (inverting) plate swing, and normalization is by the gain's magnitude. And voice::core applies the demodulator's own grid-leak inversion explicitly — a growing envelope drives that grid toward cutoff — so the two inversions put the demodulator's plate in phase with the envelope while the curve has meanwhile acted on the underside. drive now sweeps harmonic content 0.221 → 0.344, monotonically.

The gain-staging lesson from fuzz.h was applied here from the start rather than learned again: each stage is normalized by its own small-signal gain, so drive changes the distortion and not the level.

The detector is an identity, not a simplification

The plan's instruction for this stage was "synthesize the difference tone directly as a sinusoid", and catching that as a mistake is the most valuable thing this build did.

The paper's 0.03 % distortion figure and its licence to replace oscillators with a sinewave generator apply to the oscillators. The demodulator is not a mixer handing you a difference tone; it is an envelope detector, and the envelope of cos(Φ) + cos(Φ − φ) is 2|cos(φ/2)|, whose Fourier series puts H2 at −14.0 dB, H3 at −21.3 dB and H4 at −26.4 dB. Synthesizing a sinusoid would have discarded the instrument's largest single source of harmonics before any of the modelled stages ran.

What replaces the carrier is better than a simplification. For amplitudes 1 and depth, the envelope is exactly

sqrt(1 + depth² + 2·depth·cos(2π f t))

so the 80 kHz carrier drops out of the arithmetic rather than being approximated away. Running the published RC detector on that closed form — instant attack through the diode, 200 µs decay through R4·C21 — reproduces a full heterodyne-plus-diode-plus-RC simulation to within 0.10 dB on every harmonic at every pitch tried (ondes.ipynb §2).

There is one systematic difference, and it is worth knowing it is systematic rather than noise: the closed form sits a uniform 3.0–3.2 % high, because a follower chasing real carrier half-cycles never quite reaches the peak between them. On a synthesizer with a level control, that is a constant.

The detector's characteristic pitch dependence comes along free, out of the same 200 µs: H2 runs −14.0 dB at A2 to −19.3 dB at A6, and the level falls 2.0 dB across those five octaves.

And a bonus nobody planned: because the closed form is parameterized by the two oscillator amplitudes, oscillator balance becomes a physical timbre control. depth is a real mismatch between two real oscillators, not an invented knob.

Three measurements that lied, and what they were doing

All three were committed to a notebook or a header before being caught.

Too few periods. The first measurement of the detector's harmonics at low pitch used a window holding about 2.75 periods of the fundamental. Spectral leakage at that resolution dominated everything, and it produced a confident, wrong claim in the header: "−9.8 dB at A2, level falls 9.7 dB". Redone with 60 cycles, the real answer is −14.0 dB and 2.0 dB. Both numbers were in a shipped header before the recheck.

Probing where the answer is exactly zero. The aliasing scenario probed half-integer harmonics of a tone that was exactly periodic in the analysis window. Those bins are analytically zero, so it measured −281 dB and passed triumphantly. Fixed by computing the actual fold frequencies for a tone at 2637 Hz — deliberately not a submultiple of 48 kHz — and skipping folds that land near real harmonics. This is the same family of error fuzz.h records under "choosing a tone that divides the sample rate", committed again in a different disguise.

Stopping the sweep at the first plateau. The header initially claimed "4× is the knee, then flat". The notebook's own more careful run — settled state, 131072-point Hann — showed 8× continuing to improve in the top octave. Corrected to "never worse", with the full table in the header, the test comment and the notebook.

The evidence that closed an open question in fuzz.h

fuzz.h measured its oversampling sequence going the wrong way — 4× worse than 2× — and had left an untested hypothesis behind: that the culprit is imaging, since zero-stuffing by N leaves N−1 images for one filter to suppress, and residual images entering a nonlinearity intermodulate into products that are not harmonics of the input.

This file runs the same 8th-order Butterworth chain around a comparably hard nonlinearity, and its sequence never reverses:

tone1×2×4×8×
587 Hz−79.3−91.2−104.5−103.8
1175 Hz−65.8−77.2−90.6−92.5
1760 Hz−57.6−70.9−81.1−82.2
2637 Hz−51.1−61.4−71.8−83.8
3520 Hz−45.4−56.8−67.0−74.2

The difference between the two files is exactly the hypothesis: this object is a source. Nothing is zero-stuffed on the way up — the detector simply runs fast — so there are no images at all.

That was evidence, not proof — the nonlinearities differ too, and one confounded comparison does not settle a question. But it was the first evidence either way, and it pointed somewhere specific enough to act on.

Acting on it settled it. fuzz.h now cascades one 2× stage per doubling instead of zero-stuffing by N once, each stage filtering at a corner that never tightens however deep the cascade goes. Its reversal is gone — worst step-up past 2× is a ratio of 1.017 — and its 4× and 8× improved by two to four orders of magnitude, for about 5 % more CPU. This file needed no change, having no upsampler to fix.

Worth naming the shape of it, because it is not the usual one: the evidence that resolved a two-wave-old open question in one file came from building a different file that happened to differ in exactly the right variable. It was not designed as an experiment. It was noticed, written down in both headers as evidence rather than proof, and left where the next person would trip over it.

A wrapper test that found a kernel bug

tap.ondes~'s Min-level test asserts something a patcher would otherwise file as a bug report: with the key at rest, the object is exactly silent. It failed.

voice::set_smooth_ms set the voice's own ramps but never forwarded to touche::key, which keeps its own slew. So a key sitting at zero with @smooth 0 still sounded for 20 ms after every parameter touch.

This is the two-layer split working the way it is supposed to. The kernel suite tests DSP promises; the wrapper suite tests what a patcher will actually observe, and those are not the same set. The fix landed in the kernel with its own scenario, not in the wrapper.

What this file will not do

The real instrument has switchable waveform registers. Their filter shapes are in none of the sources obtained. Adding them from imagination is the one thing this file is careful not to do, and the omission is stated in the header, the object description and the reference page rather than left as a gap someone might charitably fill later.

Two controls are choices — where the intensity key sits (keyplacement) and the coupling transformer's winding sense (polarity) — because the paper's five stages do not settle either. Both are labelled as choices, and both were measured to confirm they are audible ones: about 0.09 and 0.12 of total harmonic content respectively.

Checkpoint

A stage that required no design because the paper published the model and its fitted parameters. A detector that is exact rather than approximate, and cheaper than the thing it replaces. One sign error that inverted the meaning of a distortion knob and was invisible at any single setting. Three measurements that lied in three different ways, all recorded. The evidence that closed fuzz.h's oversampler question, from a file that happened to differ in exactly the right variable and was not built as an experiment. And a wrapper test that found a kernel bug, which is the split doing its job.

How to read a recipe

The first parts of this book keep two promises: the object chapters say what each tool is for, and the machine chapters say why to trust it. This part makes a third kind of promise. A recipe puts several objects on one patch cord and chases a specific sound — a record you have heard, an instrument you have coveted — and tells you honestly how close the kit gets.

Recipes are held to the house rules, adapted:

  • Every knob named exists, spelled the way the attribute is spelled. A recipe is checkable against the reference pages; if it says @decay 0.8, that attribute takes that value on the shipping object.
  • Settings are starting points, not measurements. A recipe's numbers get you into the neighborhood; your ears walk the last block. Where a chapter number is a measurement (a decay time, an alias floor), it still cites the executed notebook or pinned test that carries it — the recipes borrow those numbers rather than re-deriving them.
  • Provenance stays honest. When a recipe chases a record, it says what is documented about how that record was made and what is folklore. When it chases an instrument, it leans on the same published analyses the kernels were built from. What a recipe never does is claim to be the record — mix, room, tape, and hands are not in the box.
  • Every recipe ranks its ingredients. The house habit from the Moog recipe in the oscillator chapter: list what each element buys, in order of importance, so you know what to cut first when CPU or taste says so.

Each recipe has the same skeleton: the sound and where it came from, the signal chain, the settings (tables for knobs, grids for patterns), what each ingredient buys, and — because every tool is sometimes the wrong tool — when to leave the recipe and cook something of your own.

One machine, four decades

The TR-808 sold poorly, was discontinued in 1983, and then spent forty years becoming the most influential drum machine ever built — not by being realistic, but by being itself in four different genres' hands. This recipe visits four of those hands: the 1982 electro of "Planet Rock," the same year's slow soul of "Sexual Healing," the tuned-kick boom of Miami bass, and the half-time rolls of trap. Same eight circuits every time; what changes is the pattern, the accents, and which knob someone dared to turn all the way up.

One honesty note before the first grid: these patterns are starting points, not transcriptions. Where a record's production story is documented, the recipe says so; the grids themselves are the versions ears agree get you into the neighborhood, and your ears finish the trip. The voice knobs, on the other hand, are exact — every attribute below is spelled as the shipping object spells it, and the calibration numbers behind the voices live in the drum machine chapter and the tr808_calibration.ipynb notebook.

The scaffold every recipe shares

One phasor~ is the transport; every tap.808.seq~ row reads it; every row's output cable is a voice's trigger input. The phasor's frequency for a 16-step bar of 4/4 is BPM ÷ 240 (four beats per cycle, four sixteenths per beat). Rows fed the same ramp are sample-locked forever — that is the sequencer's phase-derived design (see its machine chapter), and it is why nothing below mentions sync.

phasor~ (BPM/240)
   ├── tap.808.seq~  ──▶ tap.808.kick~   ──┐
   ├── tap.808.seq~  ──▶ tap.808.snare~  ──┤
   ├── tap.808.seq~  ──▶ tap.808.hat~    ──┼──▶ +~ ──▶ tap.limi~
   ├── tap.808.seq~  ──▶ (open) hat inlet 2┤
   └── tap.808.seq~  ──▶ tap.808.cowbell~──┘

Program a row with two lists: hits (which of the 16 steps sound, 1/0 per step) and accents (which sounding steps lean, 1/0 per step). An accented step emits the row's accented level (default 0.5), a plain step emits plain (default 0.01) — those defaults are the hardware's accent knob at noon, and they matter more than they look, because a voice's trigger amplitude is a voltage on the 4–14 V bus: an accented hit is punchier and differently voiced, not merely louder. Raise accented toward 1.0 when a groove should hit like the accent knob cranked. Single steps tweak with step <n> <velocity> (1-based), and each row's 16 slots (store/recall) hold your fills.

In the grids below, X is an accented hit, x a plain one, . a rest.

1982, the Bronx via Düsseldorf: the "Planet Rock" kit

The documented part: Afrika Bambaataa and producer Arthur Baker built "Planet Rock" on a rented TR-808, borrowing Kraftwerk's melodies, and its kit — dry kick, clap-snare backbeat, offbeat cowbell — became the electro sound. The orchestra stabs were a sampler's; everything percussive is the machine's.

step:    1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
kick:    X . . . . . . x . .  x  .  .  .  .  .
snare:   . . . . X . . . . .  .  .  X  .  .  .
clap:    . . . . X . . . . .  .  .  X  .  .  .
closed:  x . x . x . x . x .  x  .  x  .  x  .
open:    . . . . . . . . . .  .  .  .  .  x  .
cowbell: . . x . . . x . . .  x  .  .  .  x  .
  • Tempo ≈ 129 BPM → phasor~ 0.5375.
  • tap.808.kick~: @decay 0.35 @tone 0.55 — the electro kick is short and clicky, not the boom (that comes later in this chapter).
  • tap.808.snare~: @tone 0.6 @snappy 0.7; layer tap.808.clap~ on the same backbeat row — the clap-plus-snare composite is half the sound.
  • tap.808.cowbell~ on the offbeats, @level 0.6. The drum machine chapter's line stands: more cowbell is a patching decision.
  • Hats: closed 8ths; the open hat answering just before the bar turns.
  • Fill: store 1 the main pattern, program the classic descending-tom fill (tap.808.tom~, @size high → mid → low on three rows) into slot 2, and recall 2 a bar before the phrase ends — quantize cycle (the default) swaps it exactly on the downbeat.

1982, Ostend: the "Sexual Healing" slow jam

The documented part: Marvin Gaye programmed the TR-808 himself for "Sexual Healing," and it became one of the first major hits carried by the machine — proof in the same year as "Planet Rock" that the same circuits could whisper. The kit is soft, sparse, and riding the plain/accent distinction rather than density.

step:    1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
kick:    X . . . . . . x . .  .  .  .  .  .  .
snare:   . . . . x . . . . .  .  .  x  .  .  .
closed:  x . x . x . x . x .  x  .  x  .  x  .
open:    . . . . . . x . . .  .  .  .  .  x  .
claves:  . . x . . . . . . .  x  .  .  .  .  .
  • Tempo ≈ 94 BPM → phasor~ 0.3917.
  • tap.808.rim~ @model claves — the high tick is the hook of the kit. @level 0.5 keeps it a seasoning.
  • tap.808.kick~: @decay 0.6 @tone 0.35 — rounder than electro, still polite.
  • tap.808.snare~: @snappy 0.35 @tone 0.4 — more drum, less noise.
  • Leave the sequencer's plain level at its 0.01 default and place accents sparingly; at this tempo the difference between a 4 V hit and a half-accented one is the entire feel.
  • A touch of @swing 0.15 on the hat row loosens the grid the way a human thumb on the start button did.

Late eighties, Miami: the kick is the bassline

Miami bass turned the kick's decay knob to the top and discovered the 808's bass drum is a tuned instrument — a bridged-T resonator whose fundamental sits near 49 Hz (measured within 2.4 % of a real unit across the knob grid; see the calibration pass in the drum machine chapter). Turn decay up and it rings for seconds; give two copies two tuning ratios and you have a two-note bassline.

step:            1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
kick A (root):   X . . . . . . . . .  X  .  .  .  .  .
kick B (fourth): . . . . . . X . . .  .  .  .  X  .  .
snare:           . . . . X . . . . .  .  .  X  .  .  .
closed:          x . x x x . x x x .  x  x  x  .  x  x
  • Tempo ≈ 126 BPM → phasor~ 0.525.
  • Two tap.808.kick~ objects, two rows. tuning is a ratio of the stock fundamental, so target Hz ÷ 49 ≈ your setting: kick A @tuning 1.0 (G1, where stock already sits), kick B @tuning 1.33 (≈ C2, the fourth). Both @decay 1.0 — the whole genre is that knob at the top.
  • @tone 0.2 keeps the click out of the way of the ring; @attack 0.5 softens the punch mechanism if the notes should bloom instead of hit.
  • tap.808.snare~ @snappy 0.8 @drive 6 — the swing-VCA drive is the crack that cuts through the sub.
  • Watch the sum: two ringing kicks stack. tap.limi~ on the bus is the modern answer; riding level per voice is the period one.

The 2010s: trap, and the arithmetic of rolls

Trap keeps Miami's tuned, sustained kick and moves the snare to beat 3 — the half-time frame — then spends all its rhythmic budget on hi-hat subdivision games. Those games are where this sequencer's phase-derived design pays off: rows of different lengths off one phasor divide the same bar differently, so a 32nd-note roll row and a 16th-triplet roll row are just length 32 and length 24 — polymeter as arithmetic, measured in the sequencer notebook.

step:              1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
kick:              X . . . . . . x . .  x  .  .  .  .  .
snare:             . . . . . . . . X .  .  .  .  .  .  .
closed (len 16):   x . x . x . x . x .  x  .  .  .  .  .
roll    (len 32):  steps 25–32 hit, plain    (32nds on beat 4)
triplet (len 24):  steps 19–24 hit, plain    (16th triplets, beats 3–4)
  • Tempo ≈ 140 BPM → phasor~ 0.5833.
  • Tune the kick to the song's key with the ratio table: E1 ≈ @tuning 0.83, F1 ≈ 0.88, G1 = 1.0, A1 ≈ 1.12. @decay 1.0 @tone 0.15, and keep sigh at its default 1.0 — the pitch relaxation is the 808-bass glide everyone samples.
  • The roll rows: program hits only on their last steps (as above), leave them muted (@mute 1), and unmute for the bar that needs the roll — or keep separate patterns in slots and recall. Fast rolls do not machine-gun: the voices' filter states persist across triggers, so a roll interferes with the ringing tail like hardware (pinned by the family's tests; see the drum machine chapter).
  • Alternate hat voicing per unit: @seed is which 808 you own, and @tolerance 0.3 puts the metal bank's oscillators off-grid the way resistor variance really does (Werner et al. measured up to ~20 % — the chapter has the numbers). Two hat objects, two seeds, panned, is a stereo kit for free.

What each ingredient buys

  1. The pattern and its accents. Four decades of genre difference above is mostly the grids. The accent flags are not dynamics polish — they are the hardware's second voicing per drum. Spend your time here.
  2. decay and tuning on the kick. One knob separates electro from Miami; one ratio puts the kick in the song's key.
  3. The composite backbeat. Clap + snare on one row (electro), or snare drive (Miami, trap) — the backbeat carries the genre signature after the kick.
  4. Polymeter rows for rolls. Two extra rows, two length values, and the trap chapter of the machine's biography writes itself.
  5. seed/tolerance on the metal. Seasoning, in the salt sense: invisible until you A/B two units.

When to leave the recipe

  • You want those records, exactly. Mix, tape, room, and a human on the start button are not in the box; at some point the honest tool is the actual sample.
  • You want velocity-per-step expression. The velocities list gives a row continuous 0..1 levels — but note it trades away the two-level hardware model; the accent bus is the 808's own idiom.
  • You want 909, LinnDrum, or DMX. Different circuits, different machines — this family models one instrument, and its refusal to be generic is the point.

Three oscillators into a ladder

The oscillator chapter ends with the Moog recipe's core — the three-voice saw stack and the driven ladder — and ranks what each ingredient buys. This recipe finishes the instrument: the two envelopes, the amplifier, the gate, and the settings that turn one signal chain into the two patches everyone actually means by "Moog" — the bass that walks and the lead that sings. The model here is the classic three-oscillator monosynth voice: three oscillators into a mixer, one four-pole ladder, one loudness contour, one filter contour, glide on the pitch. Nothing below requires an object the package doesn't ship. And once the voice stands, the next recipe drives it at the records with names on them — Winwood, Worrell, Wright, Emerson.

Companion material: the oscillator chapter (the stack's rationale and the analog section's ranges), the ladder chapter (every filter number below is measured in its notebook), and the reference pages for tap.adsr~ and tap.vca~.

The voice, wired

pitch (midi note) ──▶ mtof ─┬─▶ tap.vco~ (voice 1)──┐
                            ├─▶ tap.vco~ (voice 2)──┼─▶ *~ 0.36 ─▶ tap.ladder~ ─▶ tap.vca~ ─▶ out
                            └─▶ ÷2 ─▶ tap.vco~ (3)──┘                    ▲              ▲
gate (0/1 signal) ──┬───────────▶ tap.adsr~ (filter contour) ─▶ *~ amount ─▶ +~ base ──┘│
                    └───────────▶ tap.adsr~ (loudness contour) ─────────────────────────┘
  • Pitch arrives as note frequencies (floats into each tap.vco~ left inlet; halve for voice 3's octave-down). The oscillators' own smooth ramp is the glide knob — no portamento object exists or is needed.
  • Gate is any signal that rises above tap.adsr~'s threshold and back — the envelope reads the gate by level, per sample. A tap.303.seq~ gate output (1.0 plain, 2.0 accented) drives it directly, which also gets you slides for free; so does a MIDI-driven 0/1 signal, or the trigger 1 / trigger 0 attribute messages for mouse-driven patching. The default mode analog gives the envelopes below the RC curves a Model D actually had; velocity (off by default) lets the gate's amplitude scale the hit.
  • The filter contour scales into the cutoff's signal inlet: envelope × amount (Hz) + base (Hz) into tap.ladder~'s right inlet. The classic panel's "amount" knob is your *~.
  • The loudness contour multiplies the ladder's output — tap.vca~ with the envelope into its gain inlet keeps the option of @circuit warm saturation later.

One period-correct honesty note: the original panel's contours are attack/decay/sustain with a release switch (release equals decay, on or off). tap.adsr~ gives the full four stages; set release equal to decay and you have the switch's "on" position.

The stack and the ladder

The three-voice table is the oscillator chapter's, reproduced so this page patches alone:

voicefrequencydetunedriftseed
1f−4811
2f+5822
3f ÷ 2+21033

All three: @shape 2 (saw), @jitter 3 @track 2 @imperfect 0.3, and smooth per the patch below. Sum through *~ 0.36 (≈ 1/2.8, headroom for three voices), then tap.ladder~ at the chapter's voicing: @mode lp24 @resonance 0.35 @drive 9 @asym 0.45 @comp 0.25. Keeping comp low preserves the authentic passband droop; drive 9 sits where the ladder notebook measures the tanh stages just starting to thicken (3.5 % THD at 8 dB). Spend the character budget in the filter first — the chapter's measurements are the argument.

Patch one: the bass

The left hand of a decade of records: short filter contour, no vibrato, glide short enough to read as punch rather than portamento.

controlsetting
all tap.vco~ smooth25 ms
filter tap.adsr~@attack 2 @decay 220 @sustain -18 @release 220
filter amount / base2500 Hz / 120 Hz
ladder resonance0.25
loudness tap.adsr~@attack 2 @decay 400 @sustain -3 @release 120

The sound lives in the filter contour's decay: 220 ms is the "wah" that articulates each note. Shorten toward 120 ms and it turns percussive; lengthen toward 400 ms and it turns brassy. For a rounder, more sub-friendly bass, drop drive to 3 and asym to 0.2 — the even harmonics are lovely on a lead and muddy on a bass amp. If anything downstream cares about DC, remember the ladder chapter's warning: an asymmetric saturator can leave a small signal-dependent offset — tap.dcblock~ after the VCA is one object of insurance.

Patch two: the lead

The singing version: longer glide, opened filter, resonance high enough to color but under the edge, and the release switch "on."

controlsetting
all tap.vco~ smooth80 ms
filter tap.adsr~@attack 15 @decay 600 @sustain -8 @release 600
filter amount / base4000 Hz / 300 Hz
ladder resonance0.55
loudness tap.adsr~@attack 8 @decay 300 @sustain -2 @release 350

Two moves push it from good to that sound:

  • Play the glide. 80 ms of smooth means overlapping note changes swoop; detached ones barely bend. The keyboard articulation is the vibrato.
  • Lean on the octave voice. Pull voice 3 up to f (unison) for the hollow reedy register, or leave it at f ÷ 2 and drop voice 2's level for the fat fifth-less stack. The interp-timed preset morph (store / recall <slot> <ms>) can glide between these voicings mid-phrase — a patch element the hardware never had.

What each ingredient buys

In order — and, per the house rule, cut from the bottom when CPU or taste says so:

  1. The stack. Three free-running voices at ±cents is most of the sound (the oscillator chapter's argument, with its measurements).
  2. The ladder. Drive, asymmetry, and the low-comp droop — the character budget.
  3. The filter contour. The one envelope listeners hear as "the synth's voice." Its decay is the most audible 100 ms in the patch.
  4. Glide. Free, iconic, already in the oscillator.
  5. The loudness contour. Keep it simple; the filter does the talking.
  6. The analog section. drift/jitter/imperfect at the chapter's moderate settings — salt, not sauce.

When to leave the recipe

  • You want polyphony. This is a monosynth voice; mc.-wrapping the whole chain gives you many monosynths, and a real polysynth patch wants per-voice envelopes and different discipline.
  • You want the 303 instead. The couplings that make acid are a different instrument — tap.303~ refuses to be decoupled, and that refusal is its chapter.
  • You want clean. Every stage here has an opinion — tap.svf~ and tap.fourpole~ are the polite siblings when the patch needs a filter, not a character.

The patches with names on them

The previous recipe built the three-oscillator voice. This one drives it at four records — a blue-eyed-soul hook, the bassline that retired a bass player, a singing art-rock lead, and the one-take modular solo that started it all — and, along the way, answers a fair question: if "Lucky Man" was played on a Moog modular, does the kit need a modular object?

The provenance rule from the part opener applies double here, because gear folklore is a genre of its own. For each patch the chapter says what is documented about the record and what is reconstruction. And the standing disclaimer stands: these settings chase the sound; the hands, the tape, and the mix stay on the record.

Every patch below is a delta against the wiring and tables of Three oscillators into a ladder — build that voice first. Two performance tools recur, so here they are once:

  • Vibrato is the oscillator's own now: @vibrato (depth in cents, so the musical width holds in every register), @vibrato_rate (Hz), and @vibrato_delay (ms) — the onset fades in through that time constant and re-arms on every new note, which is most of what makes a lead "sing." ±10 cents at 5.5 Hz with a few hundred milliseconds of delay is the classic setting. (This chapter's first draft had to print a scaling formula into the Hz-calibrated FM inlet here; that formula became the improvements plan's §2, and §2 became these attributes — the audit worked.)
  • Sequenced lines: tap.303.seq~ emits pitch as a MIDI-note signal and a gate at 1.0/2.0 — mtof~ turns the pitch into Hz for the oscillators' signal inlets, and the gate drives tap.adsr~ directly (it opens above 0.5). The Moog voice sequenced this way is the classic synth-line scaffold, slides included.

The Winwood hook — "While You See a Chance" (1980)

What's documented: Winwood played essentially everything on Arc of a Diver himself, synthesizers included; accounts of the rig put Moog monosynths at the center of it. The reconstruction: the opening hook is a brassy, open-filter lead with a fast attack and just enough glide to round the corners — a patch that sits between horn section and organ, which is very much a keyboardist's lead.

Deltas from the lead patch:

controlsetting
voices1 and 2 only, at f, detune −6 / +6; retire voice 3
all smooth40 ms
ladder@resonance 0.3 @drive 6 @asym 0.3
filter contour@attack 5 @decay 500 @sustain -6 @release 400, amount 4500 Hz, base 400 Hz
loudness contour@attack 5 @decay 200 @sustain -2 @release 250

The brass illusion is the filter contour's sustain sitting high (−6 dB): the filter opens and stays open, so the tone holds its brightness through the note instead of wah-ing. Play the hook in clean detached eighths — the 40 ms of glide only speaks when notes touch.

The bassline that retired a bass player — "Flash Light" (1977)

What's documented, and gloriously so: Bernie Worrell built Parliament's "Flash Light" bassline by stacking Minimoogs — the story is told with the number three attached — playing the line keyboard-style under Bootsy Collins' guitar. This is the patch where the previous chapter's "the stack is most of the sound" rule gets its funk citation.

Deltas from the bass patch:

controlsetting
voices1 and 2 at f (detune −7 / +7), voice 3 at f ÷ 2, its gain −6
all smooth35 ms
ladder@resonance 0.6 @drive 12 @asym 0.5 @comp 0.2
filter contour@attack 1 @decay 150 @sustain -24 @release 150, amount 2200 Hz, base 90 Hz
loudness contour@attack 1 @decay 250 @sustain -6 @release 100

The rubber is the filter contour: near-instant attack, short decay, and a sustain low enough (−24 dB) that every note is a squelch that immediately ducks. resonance 0.6 puts a vowel on the squelch; drive 12 into the tanh stages is the fat (the ladder chapter measures 16.5 % THD up there — that's the point). Play staccato sixteenths with octave pops; let the 35 ms glide smear only the connected passing notes. If the low end blurs, this is the one patch where comp earns its raise: 0.2 keeps some droop-era character while returning enough passband to anchor the root.

The singing lead — "Shine On You Crazy Diamond" (1975)

What's documented: Richard Wright's rig in the Wish You Were Here sessions included a Minimoog, and the singing synth lead lines in "Shine On" are credited to it. The reconstruction: a nearly clean patch — this lead's beauty is restraint, a barely-driven filter, and vibrato that arrives late.

Deltas from the lead patch:

controlsetting
voices1 and 2 at f, detune −2 / +2 — a shimmer, not a chorus
all smooth15 ms
ladder@resonance 0.15 @drive 3 @asym 0.2
filter contour@attack 30 @decay 900 @sustain -10 @release 700, amount 3000 Hz, base 250 Hz
loudness contour@attack 8 @decay 300 @sustain -2 @release 500

Then spend all your effort on the vibrato: @vibrato 10 @vibrato_rate 5.5 @vibrato_delay 400 — ten cents, arriving late, re-arming on each new note so held phrase-endings bloom while passing notes stay plain. The patch is deliberately close to the ideal oscillator — imperfect 0.2, drift at the polite end — because the expressive load is carried by the hands, and everything the analog section adds here it adds to sustained exposed notes.

The one-take solo — "Lucky Man" (1970), and the modular question

What's documented: Keith Emerson's solo on "Lucky Man" was played on his Moog modular system and famously kept from an improvised take — one of the first Moog solos on a rock record, and for a generation of listeners the first synthesizer they ever heard. The sound: a huge unison lead whose actual melodic content is mostly portamento — sweeps and dives across octaves, the glide circuit played as the instrument.

Deltas from the lead patch:

controlsetting
voicesall three; voice 3 up at f (unison), detune −5 / +4 / +7
all smooth280 ms
ladder@resonance 0.2 @drive 8 @asym 0.4
filter contour@attack 10 @decay 800 @sustain -4 @release 600, amount 5000 Hz, base 800 Hz
loudness contour@attack 10 @decay 300 @sustain -1 @release 400

At 280 ms of smooth, pitch is a place you travel to: hold a note, strike one two octaves up, and the voice draws the line between them. That is the solo. The filter stays essentially open (sustain −4 dB) because the record's drama is in pitch, not timbre.

So — does the kit need a Moog modular object? No, because you are holding one. A modular synthesizer is oscillators, filters, envelopes, and amplifiers with no fixed routing; the panel of patch cords is the product. In this package the modules are tap.vco~, tap.ladder~, tap.svf~, tap.adsr~, tap.vca~, tap.noise~, and the sequencer pair — and Max itself is the patch panel, with the routing freedom no hardwired monosynth voice (and no single "modular object") could offer. Everything Emerson's system did on that solo — voices summed to one filter, one loudness contour, glide on the pitch source — is the previous chapter's wiring diagram; what the modular added was the freedom to have wired it otherwise, and that freedom is the patching environment you are already in. The one genuinely modular idiom worth calling out is the sequenced line: tap.303.seq~ → mtof~ → the stack, gate → tap.adsr~, is the Moog-sequencer scaffold of the Berlin school and "I Feel Love"-era disco — no new object required, slides included.

What separates the four

The instructive part of putting these side by side: the signal chain never changed. What moved:

  1. The filter contour's sustain. High and it's brass (Winwood), open and it's drama (Emerson), low and it's rubber (Worrell). One attribute spans the genre map.
  2. smooth. 15 ms is articulation, 40 ms is rounding, 280 ms is the melody itself.
  3. drive and resonance. The funk patch is the only one leaning hard on both — and it's the one imitating three stacked instruments.
  4. The hands. Delayed vibrato, staccato versus legato, when not to play — the parts of the record the recipe honestly can't ship.

When to leave the recipe

  • You want the record's whole arrangement. The hook was never alone: Winwood's is doubled, Worrell's sits under a live band, Wright's floats on tape-delayed guitars. The patch is the voice, not the mix.
  • You want polysynth-era sounds. Prophets and Oberheims are a different architecture — per-voice envelopes on real polyphony — and imitating them with mc. stacks of this voice flatters neither.
  • You want the sequenced-modular genre. Start from the scaffold above, but that recipe deserves its own chapter — it lives in the plan file's backlog with "I Feel Love" written on it.

Move a knob while it loops

The documented origin story of acid house is an instruction manual for this recipe: in Chicago around 1985–87, Phuture (DJ Pierre, Spanky, Herk) let a secondhand TB-303 loop a pattern and turned the knobs while it played — "Acid Tracks" is twelve minutes of that. The lesson generalizes: an acid line is not a melody with a sound; it is a loop plus a hand. The pattern's job is to give the couplings something to chew on — accents for the bloom, slides for the vowels — and the performance is cutoff, resonance, and envmod moving in real time.

Everything measured here is borrowed from the acid machine chapter and its notebooks (tb303.ipynb, step_seq.ipynb).

The scaffold

phasor~ (BPM/240) ──▶ tap.303.seq~ ──pitch──▶ tap.303~ ──▶ out
                                  └──gate───▶   (right inlet)

One bar of 16 steps per phasor cycle (BPM ÷ 240, the drum scaffold's math); ~125 BPM → phasor~ 0.5208. The sequencer's pitch and gate outlets are the voice's own contract — accents ride the gate at 2.0, slides are pitch changes under a held gate, so the voice's ~60 ms RC does the glide.

A line to start from

Program per step (step <n> <pitch> [accent] [slide], rest <n>) or per lane. A serviceable opener in A — and, as everywhere in this part, a starting point, not a transcription:

step:    1  2  3  4  5  6  7  8  9  10 11 12 13 14 15 16
pitch:   33 33 45 33 33 36 33 31 33 33 45 47 33 33 31 36
gate:    x  x  x  x  .  x  x  x  x  x  x  x  .  x  x  x
accent:  A  .  .  A  .  .  A  .  .  A  .  .  .  .  A  .
slide:   .  .  .  .  .  .  .  S  .  .  .  S  .  .  .  .
pitches 33 33 45 33 33 36 33 31 33 33 45 47 33 33 31 36
gates   1 1 1 1 0 1 1 1 1 1 1 1 0 1 1 1
accents 1 0 0 1 0 0 1 0 0 1 0 0 0 0 1 0
slides  0 0 0 0 0 0 0 1 0 0 0 1 0 0 0 0

The ingredients that make it acid rather than bass: the octave jumps (33 → 45), at least one slide into a note (the flag sits on the target step), rests that let the filter close, and accents placed where the groove leans — not where the melody peaks.

The voice

recall 1 is the factory "squelch" and a fine start. Explicitly:

@waveform saw @cutoff 500 @resonance 0.9 @envmod 0.7 @decay 300 @accent 0.8

Then the moves, in the order a set builds:

  1. Ride cutoff. 300 → 2000 Hz over eight bars and back. This is the genre. Remember the modeled envmod law: 2/3 of the envelope's sweep sits above the knob, 1/3 below, and the resting point shifts as you turn it — the knobs interact like the hardware because the interaction is modeled.
  2. Stack the accents. Runs of accented notes at high resonance are the wow: the C13 capacitor doesn't fully discharge between close accents, and the measured cutoff-peak bloom across a run is ×1.94. Put three accents in a row somewhere and listen to the third one open.
  3. Raise resonance into the break. Past 1.0 is the documented bend territory — a stock 303 never self-oscillates, and neither does this filter until you push it there deliberately.
  4. waveform square for the hollow verse, saw for the drop.
  5. vca warm thickens exactly where the hardware does — measured 5.4 % difference signal on quiet notes, 11.5 % on hot accents.
  6. Transpose, don't re-program: transpose -12 on the sequencer for the sub-drop, +5/+7 for the question-answer sections. It shifts live, like the hardware's transpose mode without the mode.

After the voice

Acid techno's other instrument is the distortion pedal: tap.overdrive~ after the voice, driven hard, is the documented lineage (a 303 into a screaming feedback overdrive is half the harder end of the genre). Keep mute in reach on the sequencer for breakdowns — it drops the gate but the clock keeps running, so the line re-enters exactly in place.

When to leave the recipe

  • You're programming melodies. If the line only sounds right without slides or accents, it isn't an acid line yet — or it wants the generic bass rig (tap.vco~ + tap.svf~ + tap.adsr~) instead of this voice's refusals.
  • You want the filter alone — tap.diode~ gives you the 303's ladder on any source, squelch and all, without the biography.
  • You want hands-free evolution. The 303 rewards a hand on a knob; if the patch must run itself, store extremes in the voice's preset slots and ride timed recall morphs instead — a different instrument, honestly.

The ostinato machine

Two documented lineages share one patch. The Berlin school — Tangerine Dream's Phaedra (1974) above all — put a Moog modular's step sequencer on stage and let a filtered ostinato run for twenty minutes while hands moved the cutoff. Three years later Giorgio Moroder and Donna Summer's "I Feel Love" built an entire hit from a sequenced Moog modular bassline. The Recipes part's Moog chapter argued you already own the modular — Max is the patch panel; this recipe is that argument cashed in: the sequencer pair driving the three-oscillator voice.

The scaffold

phasor~ (BPM/240) ──▶ tap.303.seq~ ──pitch──▶ mtof~ ──▶ slide~ ─┬─▶ tap.vco~ ──┐
                                  │                             └─▶ tap.vco~ ──┼─▶ *~ ─▶ tap.ladder~ ─▶ tap.vca~
                                  └──gate──┬─▶ tap.adsr~ (filter) ─▶ *~ amount ─▶ +~ base ──▶ ▲ (cutoff)
                                           └─▶ tap.adsr~ (loudness) ────────────────────────────▶ ▲ (gain)
  • tap.303.seq~'s pitch outlet is a MIDI-note signal; mtof~ turns it into Hz for the oscillators' signal inlets.
  • Its gate outlet (1.0 plain, 2.0 accented) drives both tap.adsr~ contours directly — the envelope opens above 0.5.
  • One honest wrinkle: tap.vco~'s frequency signal inlet bypasses the smooth ramp by design ("you are the smoothing") — so sequenced pitch steps land as hard steps, and slide flags in the pattern won't glide on their own. Put a one-pole slew (Max's slide~, or rampsmooth~) between mtof~ and the oscillators; the 303's ~60 ms RC is the reference feel. Short slew = articulation, long slew = the Berlin swoop.
  • The oscillator stack, ladder voicing, and envelope tables come from Three oscillators into a ladder — the bass patch is the right starting point. One voice instead of three is period-correct for the sequenced genre and cheaper; add the stack when the line is the whole arrangement.

The line

The genre's cell is small and the sequencer's phase math does the rest (one bar per phasor cycle; a length 8 row divides it into eighths — polymeter as arithmetic, per the sequencer chapter).

The octave bounce, "I Feel Love"-school — length 8, every step gated:

step:    1  2  3  4  5  6  7  8
pitch:   33 45 33 45 33 45 33 45
pitches 33 45 33 45 33 45 33 45
gates   1 1 1 1 1 1 1 1

The Berlin cell — length 16, a contour that repeats but doesn't resolve:

pitches 33 33 40 36 33 43 36 40 33 33 40 36 31 43 36 38
gates   1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1

Then the two moves that carry twenty minutes:

  1. Transpose, sparsely. transpose 0 → 3 → 5 → 0 at phrase boundaries is the harmonic language — one message, and the armed pattern semantics keep everything on the grid.
  2. Ride the filter, slowly. The loudness contour stays short and percussive; your hand (or a very slow LFO into the +~ base) opens the ladder over minutes, not bars. The ostinato doesn't change; the light on it does.

Settings that read as the genre

controlsetting
loudness tap.adsr~@attack 2 @decay 180 @sustain -12 @release 120
filter tap.adsr~@attack 2 @decay 160 @sustain -20 @release 160, amount 1800 Hz, base 150 Hz
ladder@mode lp24 @resonance 0.3 @drive 6
slew (slide~)short; raise it only for deliberate swoops
seq@swing 0 — the genre is a grid, and the delay does the humanizing

Two period tricks worth their lines: pan alternate notes (a length 8 row of accents driving tap.pan~ recreates the famous ping-pong doubling), and put an eighth-note tap.delay~ after the voice (@feedback 40 @mix 30) — the echo, not the sequencer, is where these records' motion lives. Since its rebuild the delay interpolates (Hermite) and regenerates through a DC-blocked loop; @interp 0 remains the bit-faithful legacy mode.

Glue: an 808 closed-hat row in 16ths from the drum scaffold, mixed low. Accents land in this scaffold too: turn up tap.adsr~'s velocity sensitivity and the sequencer's 2.0-amplitude accented gates hit the envelopes harder — the loudness contour for punch, the filter contour for the quack, or both.

When to leave the recipe

  • You want the 303's couplings — slides that bloom, accents that squelch. That's the acid recipe; this scaffold trades the couplings for a filter and envelopes you choose.
  • You want generative movement. This sequencer is deliberately deterministic; probability and ratchets are future emitters, and randomness belongs to objects that own a seed.
  • The line wants to be a song. Sixty-four steps is the ceiling; past that you're composing, and a piano roll is kinder than sixty-four step messages.

The robot on the radio

The vocoder chapter closes on "the casting is everything," and this recipe is the casting call. The sound has a documented pedigree — Bell Labs speech-compression research became, in musicians' hands, Kraftwerk's robot choirs and ELO's talking skies, and the machine has never left the radio since — but the records disagree on gear and agree on craft: a bright, busy carrier, an articulate modulator, and somebody enunciating like they mean it. All three are patching decisions.

Everything structural below is pinned by the kernel's tests and explained in the vocoder chapter: 24 bands, 50 Hz–12 kHz, everything you hear is carrier.

The carrier, built properly

The eternal failure is a dull carrier — high bands with nothing in them, consonants gone. Build it in three layers:

tap.vco~ (saw, f)        ──┐
tap.vco~ (saw, f, +7 c)  ──┼─▶ +~ ──▶ tap.vocoder~ right inlet
tap.noise~ (white) ─ *~ 0.1┘
  • Two saws, a few cents apart (@shape 2, detune ±4–7): harmonics to the top of the range, and the beating keeps long vowels alive. Use the Moog recipe's stack values; skip the octave-down voice — vocoded speech reads clearest with the energy above the fundamental.
  • The s and t budget. @sibilance 0.3 is the built-in version — a seeded noise source in the top bands' carrier, gated by the modulator's high-band envelopes, arriving exactly when consonants do. The manual alternative (a tenth of tap.noise~ summed into the carrier) remains the craftier option when you want to choose the noise color yourself.
  • Pitch is the performance. The vocoder never changes the carrier's pitch, so the carrier's notes are the melody. Held chords (an mc. stack of carriers) make the robot a choir; a single line makes it a lead vocalist.

The modulator, cast against type

Articulation beats fidelity — the chapter's measured point is that band envelopes carry everything, so contrast between bands is what you feed it. A cheap dynamic mic is fine; compression helps; and over-enunciating helps more than any knob. Keep the modulator out of the mix — the machine uses it, nobody should hear it.

The three settings

patchqresponse_intervalthe craft
the talking synth20 (default)30speak in rhythm; consonants land like drum hits
the choir10–15250sing sustained vowels; the carrier chord is the harmony
rhythm transfer25–4010–20drum loop as modulator; any sustained pad as carrier

q trades crispness against smoothness (narrow separates consonants, wide blends vowels); response_interval is the mouth's speed — attack and release in one knob. gain is linear makeup, and you will need some: a band-multiplied signal lands quieter than either input.

Two wiring facts that account for most dead patches: the modulator is the left inlet (a synth weakly filtered by your voice means the cables are backwards), and a silent carrier is silence no matter how loudly you speak — pinned by test, and the fastest debugging question in vocoding.

The songbook

The famous "vocoder songs" are the best syllabus for the craft — partly because several of them aren't vocoders, and knowing which is which teaches more than any preset. Provenance below follows the part's rule: documented where it's documented, labeled reconstruction where it isn't.

"In the Air Tonight" (1981) — the ghost choir

What's documented: Phil Collins ran the verse vocal through a Roland VP-330 — a soft vocoder, voiced like a string machine, mixed under a nearly whispered dry vocal. The reconstruction: this is the anti-robot patch. Carrier: two saws, detune ±4, imperfect 0.3, no noise layer — sibilance is what you don't want here — through tap.svf~ (@type lowpass @frequency 4000) to take the glass off. Vocoder: q 8–12, response_interval 120 — wide bands and a slow mouth blur the consonants into breath. Mix the vocoded return under the dry voice, a shadow rather than a double. The dry whisper carries the words; the vocoder carries the dread.

"Mr. Blue Sky" / "Mr. Roboto" / "Intergalactic" — the front-and-center robot

ELO (EMS vocoders, documented), Styx, and the Beastie Boys are the talking-synth patch played as a lead: bright carrier, crisp bands (q 20–30), fast mouth (response_interval 15–30), and the melody in the carrier's held notes while the words ride the modulator. Kraftwerk — the genre's founders, on custom and commercial hardware across the years — sit here too, usually with a single unison line rather than chords: the robot speaks in monophony. Enunciate. Then enunciate more.

One label to keep straight: Zapp, Roger Troutman, and the P-Funk talkbox records are not vocoders — a talkbox pipes the carrier into the performer's actual mouth and the room mic hears real articulation. Chasing that sound with this object gets you a cousin, not the thing.

"Hide and Seek" (2005) — the one that isn't a vocoder

What's documented: Imogen Heap sang into a harmonizer (the DigiTech Vocalist lineage), a keyboard choosing the chord — so every sound on the record is her actual voice, pitch-shifted into harmony, breath and formants intact. That's why it doesn't sound like a robot; there is no carrier. Three routes, honestly ranked:

  1. The right tool — tap.harmony~. This record's mechanism is exactly what the object does: formant-preserving voices holding a chord over the aligned dry voice. Its recipe — with the Bon Iver patches that extend the lineage — is A choir of one. (This object exists because this section's first draft had to work around its absence; the audit worked.)
  2. The manual fallback — a shifter stack. Voice into parallel tap.shift~ objects at chord intervals (tap.semitone2ratio feeds their ratio inlets). Keep the voicing within ±7 st — plain granular shifting moves formants with the pitch, and wide intervals go chipmunk where the formant-corrected routes don't.
  3. The vibe — the choir patch above. Speak-sing into the choir row's settings with an mc. carrier holding the chords. It will sound like a vocoder doing Imogen Heap, which is its own valid sound — just don't mistake it for the record's mechanism.

The plugin-era default — an Orange-school carrier

The late-90s software vocoders (the Orange Vocoder the most loved of them) changed the default sound of the effect: where hardware vocoders leaned on whatever synth was nearby, the plugins shipped with a built-in, very bright virtual-analog carrier — so "the plugin sound" is really a carrier voicing: wide, glossy, present. One honest line first: that plugin is still a shipping commercial product, and the house provenance rule applies — nothing here reverse-engineers it. What follows is our bright VA carrier in that school, built from this package's own oscillator:

tap.vco~ (saw, f,   detune -6, seed 11) ──┐
tap.vco~ (saw, f,   detune +6, seed 22) ──┼─▶ +~ ─▶ tap.svf~ (highshelf) ─▶ carrier
tap.vco~ (saw, f+12, gain -6,  seed 33) ──┤
tap.noise~ (white) ─ *~ 0.08 ─────────────┘

All three oscillators @shape 2 @imperfect 0.2 @jitter 2; the octave-up voice adds the gloss the era is remembered for; tap.svf~ @type highshelf @frequency 6000 @gain 4 is the sheen. Vocoder settings: q 25, response_interval 20. Play the carrier in fifths and octaves rather than full triads — the brightness supplies the width, and triads in a bright carrier smear the consonant bands.

When to leave the recipe

  • You want tuned speech, not a played carrier — the corrector (tap.tune~) moves the voice itself; the vocoder wears the voice over something else. Different identity theft.
  • You want formant-shifted or gender-shifted voice — that's spectral surgery, not band gating; the corridor starts at tap.spectra~.
  • You want intelligibility above all. Twenty-four analog-style bands are a voice, not a spectrograph; if every syllable must survive, dry speech mixed under the vocoded double is the radio trick that always works.

A choir of one

This chapter exists because this book's own audit demanded it. The vocoder songbook had to label "Hide and Seek" honestly — a harmonizer, not a vocoder: pitch-shifted copies of the actual voice, formants intact, no carrier anywhere — and the package had no object for that mechanism. Now it does. tap.harmony~ holds up to four formant-preserving voices at intervals you set in semitones, over a dry path the kernel delays into alignment so chords land as chords. This recipe is how to sing through it, and its worked examples are the modern masters of the effect: Bon Iver.

The claims behind the object are measured in the executed verification notebook and pinned in the kernel's test battery (tests/harmonizer_test.cpp): across two octaves of voicings every interval lands within 0.04 cents of its equal-tempered target under the DspTap yin oracle; the dry path is sample-aligned with the voices to a 3.7×10⁻⁸ residual — why chords land as chords, not flams; a synthetic formant bump stays near home only when formant is on (band centroid 1058 → 1154 Hz on a +7 shift, versus 1439 riding the full ratio with it off); and interval glides walk the pitch through the middle instead of jumping. The engine per voice is the DspTap phase vocoder — the same peak-locked shifting and LPC formant machinery the pitch machine chapter derives.

The instrument

voice ──▶ tap.tune~ (@speed 0, key of the song) ──▶ tap.harmony~ ──▶ out

The corrector upstream is optional but it is the modern sound: hard-snap the lead first and every harmonizer voice inherits the quantized pitch, so the stack locks like a keyboard instrument instead of drifting like a choir. Skip it and the stack breathes with your intonation — older, warmer, more Crosby-Stills than Vernon.

Three controls do the character work:

  • formant on is the entire point: an octave-down voice stays you, bigger. Off is the chipmunk-chorus bend — useful, but it stops being a choir.
  • chord is the performance surface: chord -12 3 7 sets three intervals and silences the fourth voice in one message. Wire a Max chord-to-intervals mapping (played notes minus the sung root) and the keyboard chooses the harmony live — the rig the credits of the records below describe.
  • glide at the 10 ms default snaps chord changes; at 300–500 ms the stack slides between chords, which no group of human singers can do and is worth featuring, not hiding.

One honest number: latency is one FFT frame (fftsize, default 1024 samples ≈ 21 ms at 48 kHz), dry path included. For live monitoring that is audible as a slight remove — performers adapt in minutes, but mix the monitor wet so they hear the instrument, not the delay ghost.

"Woods" (2009) — the stacked chapel

What's documented: Justin Vernon built "Woods" from many overdubbed a cappella takes, each hard-tuned — a chapel of his own voice, later the foundation of Kanye West's "Lost in the World." The record's mechanism is overdubs, and the recipe respects that:

  • The live approximation: tap.tune~ @speed 0 → tap.harmony~ with chord 3 7 12, @dry 1, all through a long dark reverb (the wash settings work). One pass, four-voice chapel.
  • The faithful version: record takes — sing each chord tone through the corrector alone, layer them, and use tap.harmony~ per take only to widen (chord 12 at @level1 0.4). Stacked takes decorrelate the way overdubs do; one harmonizer pass, however good, is one performance. The difference is the difference between a choir and a string patch.

"715 - CRΞΞKS" (2016) — the Messina

What's documented: the 22, A Million credits name "the Messina," the rig Chris Messina and Vernon built to pitch-stack his live voice into chords (the Prismizer-school effect associated with Francis and the Lights, and heard on Chance the Rapper's gospel records). "715 - CRΞΞKS" is that instrument a cappella: every sound is the processed voice.

voice ─▶ tap.tune~ (@speed 0) ─▶ tap.harmony~ @dry 1 @formant 1 @glide 10
                                     chord -12 3 7    (verse color)
                                     chord -12 4 7    (the lift)
                                     chord -5 3 10    (the ache)
  • @dry 1 — the lead lives inside the stack, equal citizen, exactly what makes the sound read as one multiplied person rather than lead-plus-backers.
  • Chords change per phrase, not per note: bind each chord list to a key or a pedal and play the harmony like slow organ stops.
  • The low voice carries the weight: -12 under a falsetto lead is the record's signature register trick. Keep it at full level; thin the upper voices (@level3 0.7) when the stack gets glassy.
  • No reverb, or almost none — the record's intimacy is the dry stack right against the microphone. Resist the wash this once.

The craft notes

  1. Feed it one voice. The formant model and the intervals both assume monophonic input — the kernel's header says so, and a strummed guitar through a "choir" proves it right within two bars.
  2. Mind the sum. Dry plus four unity voices is five voices; tap.limi~ or a *~ 0.5 after the object is the standing advice.
  3. Close voicings beat wide ones. ±12 is the working span; the engine's contract runs to ±24 and the top octave of that range is a sound effect, not a singer.
  4. The corrector's speed is the era dial. 0 ms is 2016; 40 ms is 1970s session stack; bypassed is a folk group.

When to leave the recipe

  • You want the robot. No carrier here, no bands — that's the vocoder, and the two chapters together are the voice-processing fork in the road: wear the voice over a synth, or multiply the voice itself.
  • You want real ensemble. Overdub real takes; the "Woods" section's faithful version is the honest ceiling of one person's choir.
  • You want harmony that follows chords you sing. The object holds intervals; it doesn't do music theory. The keyboard (or your patch's chord logic) is the brain — which is exactly how the famous rigs work.

The staircase and the wash

Shimmer has a documented birthplace: Brian Eno and Daniel Lanois in the early eighties, feeding a pitch shifter and a reverb into each other until a guitar came out sounding like weather. The pitchaccum chapter tells the half of the story that lives inside one object — the transposer-delay loop where every pass climbs again — and its recipes sketch the pairing. This recipe is the whole patch: the spiral, the wash, and the mix decisions that keep ten seconds of accumulated fifths from eating a track.

Measured claims are borrowed from the spiral staircase (the +7-becomes-+14 accumulation, the constant-sum grain envelopes, the 0.99 feedback cap) and borrowed rooms.

The chain

source ──▶ tap.pitchaccum~ ──▶ reverb (tap.verb~ or tap.convolve~) ──▶ return
   └────────────────── dry path ───────────────────────────────────────▶ out

Run it as a send: the source stays dry and full-size in the mix, and the shimmer return comes up underneath it like backlight. On the send, tap.pitchaccum~ at mix 100 (its own dry path stays home) and the reverb wet-only.

The spiral

controlsetting
trans1 / delay1 / fb1 / gain1+12 st / 400 ms / 75 / 50
trans2 / delay2 / fb2 / gain2+7 st / 650 ms / 60 / 50
xfade60 — smooth flanks, soft attacks
modfreq / moddepth / modphase0.3 Hz / 0.1 st / 90°
followoff for chords and pads; on for monophonic lines

The two shadows are doing different jobs: the octave climbs politely (+12, +24, +36 — always consonant), while the fifth rotates the harmony (+7, +14, +21 — a fifth, then a ninth, then a #11) and is where the Eno-school mystery comes from. Pull fb2 down toward 40 when the source is already harmonically rich; push fb1 toward 90 for the endless version — the loop is capped and DC-blocked, so "too long" is an aesthetic problem, not a stability one. The touch of modulation (moddepth 0.1, with modphase 90 breathing the shadows against each other) keeps a long spiral from sounding cloned — depth stays subtle or the climb turns seasick.

The wash

Either reverb works; they fail differently:

  • tap.verb~ (the designed tail): @mix 100 @decay 8 @damping 4000 @lowpass 8000 @delay 60 @modfreq 0.2 @moddepth 0.3. The damping matters more than the length — shimmer's accumulated highs need somewhere soft to land, and 4 kHz of loop damping is the difference between glow and glass dust.
  • tap.convolve~ (the borrowed room): a long church at @mix 100 @predelay 20, and pick the IR by its top end — audition the tail alone and reject anything that rings metallic up high, because the spiral will find it. (The field guide to rooms has the audition drill.)

Order matters and is worth an experiment: spiral → reverb (above) washes the staircase — the classic. Reverb → spiral transposes the wash itself and is wilder and less controllable; the historical chains did both, depending on the record.

Variants

  • The descent: trans1 -5, trans2 -12, long delays, feedback ~50 — the staircase into the basement. Darker damping (2–3 kHz); the low accumulation muddies fast, so shorter reverb.
  • The micro-halo: trans1 +0.15, trans2 -0.15, delays 60/90 ms, feedback ~50, xfade wide, modest reverb — no spiral at all, an expensive-sounding widener that flatters pads.
  • The gesture: store the halo in slot 1 and the full +12/+7 spiral in slot 2, then recall 2 8000 as the chorus lands — the morph engine glides every parameter, and the bloom is the production moment.

When to leave the recipe

  • The mix is dense. Shimmer is backlight; on a busy arrangement it reads as mud. It earns its keep on sparse sources — one guitar, one voice, one held pad.
  • You want rhythmic echoes climbing in pitch. The delays here serve the loop, not the grid; that patch is a tempo-synced delay into tap.shift~, built by hand.
  • You want the pitch to stay put. Then it's just reverb — go straight to the field guide.

Sixteenths into a listening filter

The envelope filter earned its place in funk on documented records — the Mu-Tron-era clavinets and basses of the seventies, Stevie Wonder's "Higher Ground" chief among them — and the autowah chapter is honest about what this object is instead: a model of the Snow White AutoWah, a different, throatier circuit. You are not summoning a Mu-Tron; you are plugging into a very good pedal that listens the same way. The funk is in what it listens to — which makes this the one recipe where the settings table is half the story and your right hand is the other half.

Measured behavior cited below (the sweep law exact to the design, the RC release, the 250 Hz → ~2.5 kHz hardware span) lives in the pedal that listens and its validation notebook.

How to think about the knobs

Two of them are calibration, one is the personality:

  • sensitivity matches the pedal to your source's level and your touch. Tune it so your normal hits open the filter halfway and your hard hits open it fully — the tanh knee compresses beyond that instead of slamming. Too high and everything pins; too low and the filter ignores you.
  • bias and range set where the sweep lives: resting frequency and octaves above it. The defaults (250 Hz, 3.3 octaves) are the hardware.
  • decay is the personality: how fast the filter falls back. Tens of ms is a wah articulation on every note — the funk setting. Hundreds is a swell that rides phrasing.

The patches

patchsettings
the clav chop@sensitivity 3 @attack 2 @decay 80 @bias 250 @range 3.3 @resonance 0.7 @mix 100
the bass quackrecall 2, then @decay 150 @resonance 0.6
the slow swellrecall 3, or @decay 900 @range 2.5 on pads
the cocked wahrecall 4 — sensitivity at −60 is the envelope off; park bias at 800–1200 Hz
  • The clav chop wants sixteenth-note playing with deliberate dynamic contrast — the filter turns your accents into vowels. mode 1 (bandpass, the circuit's other tap) is quackier and noticeably quieter; make it up with gain.
  • The bass quack starts from the factory bass voicing (slot 2 — lower bias, tighter range, the GB pedal's instrument switch as a preset). Fingers, not pick, and let notes ring — the release is a real RC discharge (measured: a pure exponential, σ = 0.004) and it sounds like circuitry when you leave it room.
  • The cocked wah is the secret mode: a fixed resonant filter with bias as a manual sweep — the parked-pedal midrange honk, and slot 4 ships it.
  • direction 1 sweeps down from bias — the extension the pedal never had; reverse-envelope funk on a clean chop is startling.

The two patch points nobody uses enough

  • The sidechain (right inlet): one sound wahs another. Kick → sidechain, pad → filter is the classic; a tap.808.seq~ row (through @pulse widened impulses) makes the filter sequenced while the pad sustains — an envelope filter with a drummer's timing.
  • The envelope outlet (right outlet, 0..1 signal): the detector as a free modulation source. Scale it into tap.vco~'s FM inlet, a tap.vca~ gain, or a second filter — one performance, many destinations. (In bypass the outlet goes to zero, so tap it from a live instance.)

When to leave the recipe

  • Your source has no dynamics. A static pad through an auto-wah is a static filter — feed the sidechain something rhythmic, or use tap.svf~ with an LFO and own the motion yourself.
  • You want the filter on a knob. That's the cocked wah until you want morphing responses — then tap.svf~'s morph is the tool.
  • You want the exact Mu-Tron quack. Raise resonance, try mode 1, and know the chapter's warning stands: you're modding a Snow White. The hardware A/B pass — the notebook cell waiting for the real pedal — will tell us precisely how far the model is from its own hardware, not from someone else's.

A field guide to rooms

The convolution chapter makes one promise that changes how you shop: tap.convolve~ is exact — measured to 10⁻¹² against direct convolution — so the engine contributes no character at all. Everything the effect sounds like is the impulse response you load. That turns "how do I get a good reverb?" into "how do I find, judge, and place a good room?" — a curation problem, and this recipe is the field guide.

Companion material: the convolution chapter and its verification notebook; every measured number cited below lives in one of them.

The shopping list

An IR is a recording of a space answering a click, and the internet holds decades of them — university acoustics archives, church-recording projects, hardware-unit captures released by their communities. What to bring home, by job:

  • A church or concert hall (2–5 s). The default "make it beautiful" space. Long tails flatter sustained, sparse material and drown busy mixes — the classic trade.
  • A plate. Not a room at all — a steel sheet's dense, fast-building wash. The vocal reverb of half the records you know; sits in a mix better than any hall because it has no early-reflection "walls" to argue with the stereo image.
  • A spring. The lo-fi twang of amp reverb; gloriously wrong on drums.
  • A small real room (0.3–0.8 s). The most useful and least glamorous purchase: drums and guitars recorded dry come alive with a believable space that reads as "a room," not "an effect."
  • Not a room. The chapter's point stands in practice: any filter you can record is loadable. A guitar body IR makes a piezo pickup sound like wood; a vowel is a formant filter; a single click is a delay.

Prefer 4-channel captures when offered: the engine runs the full true-stereo matrix (LL/LR/RL/RR), and the cross-feed paths are where "being in the room" lives — measured in the notebook at exact path gains with zero leakage. A 2-channel IR runs as dual mono (no cross-feed); a mono IR is the same room on both sides.

Judging a room in sixty seconds

Load it into the buffer~, then:

  1. Send a click through and listen to the tail alone (mix fully wet). You are auditioning the IR itself — the engine adds nothing. A good tail decays smoothly darker; a flutter or a metallic ring here will be on everything you send.
  2. Check the onset. Silence before the direct sound is pre-delay baked into the capture — trim it in an editor or accept it, but know it's there, because it stacks with the predelay you set and the engine's own blocksize samples of latency.
  3. A/B at matched loudness. normalize 1 is on by default and is energy-based, so a quiet cathedral capture and a hot plate land at comparable levels — judge the room, not the gain staging.

Placing the room in a patch

  • Send, don't insert. One tap.convolve~ fed by a send bus serves the whole patch, glues sources into one space, and keeps the option of riding the send. Keep mix 100 (wet-only) on a send; use mix as an insert dry/wet only on a single source.
  • predelay before you EQ. 10–30 ms separates the dry attack from the wash and buys clarity for free — the chapter's advice, and the first knob to reach for when a reverb "swallows" a vocal.
  • blocksize by role. Live input through the reverb: 64–128 (1.3–2.7 ms at 48 kHz — the measured cost is exactly blocksize samples). Mix-bus send: 512–2048, the CPU-cheap end, where the latency reads as a little extra pre-delay you set once and forget.
  • Swap rooms as a performance move. IR swaps are atomic and click-free (measured RMS across the swap: 21.9 → 22.1) — load verse-room and chorus-hall into two buffer~ objects and rebind with set <buffer-name> on the downbeat. (The buffer is the only way in: the object binds the buffer~ named by its first argument, and re-loading a file into that buffer re-transforms the IR automatically.)

When to leave the recipe

  • You want to design the tail — decay and damping knobs, modulation, gated endings. A static IR can't; tap.verb~ is the algorithmic sibling built for exactly that.
  • You want shimmer. The wash is only half of it — the spiral half is tap.pitchaccum~, and that pairing has its own recipe in this part.
  • You want zero latency. blocksize samples, full stop; at 64 that's small, not zero.

Chords with no keyboard

The comb-bank chapter ends on "strings, chords, drones, and gestures; no guitar required" — this recipe supplies the chords. tap.5comb~'s five voices tune in Hz (freq1..5), which means voicings are numbers you can keep, trade, and morph between; below is a small book of them, plus the excitation and morph craft that turns a filter bank into an instrument.

The mechanics cited here — ring time on a log map (20 ms–100 s), Hermite tuning, warp's stretched partials, phase's midpoint pluck — are measured and explained in five strings, no guitar and its machine chapter.

Voicings to keep

Tunings in Hz; MIDI equivalents in parentheses for orientation. The notes message tunes a voicing in one gesture — up to five MIDI note numbers, fractional allowed, so just-intonation intervals land exactly (notes 45 52 57 60.86 64 is the major glow with its true 5/4 third at 275 Hz) — and the Hz attributes remain for exact ratios like the bell plate.

voicingfreq1..5character
the factory chord80 / 120 / 160 / 200 / 102the legacy preset: a root-fifth-octave stack with a rub (102 against 80)
the open fifth55 / 110 / 165 / 220 / 330 (A1, A2, E3, A3, E4)power-chord drone; nothing to clash with any source
the major glow110 / 165 / 220 / 275 / 330 (A2, E3, A3, ~C#4, E4)just-intonation major: 275 is a pure 5/4 third, warmer than 12-TET's 277.2
the dark cluster65.4 / 77.8 / 98 / 130.8 / 196 (C2, D#2, G2, C3, G3)minor with a low rub; film-cue territory
the bell plate210 / 297 / 420 / 594 / 841non-octave (√2 ratios): inharmonic, gong-ish before warp even arrives

Masters make voicings performable: freq (0..2) transposes the whole bank proportionally — chords stay chords under the glide — and res/lp scale ring and brightness bank-wide.

Ringing them

  • Drone: res1..5 85, lp toward 5000. Feed it anything quiet and sustained — pink noise at low level, a field recording, your room tone. At res ≈ 100 the bank sustains essentially forever; the input stops being audio and becomes bowing pressure.
  • Pluck: res1..5 around 60–70 and excite with clicks or a sparse tap.808.rim~ (@model claves) pattern — every tick strums the chord. Drums work; speech works eerily well (the chapter's "resonator chord").
  • Strings, stiffened: warp 40 stretches the upper partials sharp — piano-ish, then bell-ish — while the compensated main tap keeps the pitch put. Pair with lp near 3000 for the felt-hammer version.
  • The midpoint pluck: phase 100 cancels the even harmonics — the hollow, clarinet-adjacent voicing of a string plucked exactly at its middle. On the bell plate tuning it turns purely ceremonial.

Watch the sum: five ringing combs stack like five strings. Ride gain down as res goes up, and tap.limi~ on the output is cheap insurance for the res 100 lifestyle.

The gesture

The bank's real instrument is the morph engine. Store the major glow in slot 1 and the dark cluster in slot 2 — then recall 2 8000 and every frequency, ring time, and damping glides for eight seconds through tunings you never chose, Hermite interpolation keeping the sweep continuous instead of zippered. The chapter's advice stands: automate nothing else. One long morph over a static source is a complete piece of sound design; grabbing a single fader mid-morph overrides just that parameter, which is the escape hatch when the in-between territory finds something worth keeping.

When to leave the recipe

  • One resonance, surgically placed: tap.comb~ is the single unit, or tap.svf~ @type bell when you want EQ, not a string.
  • You want echoes. Combs long enough to hear as repeats are delays wearing a costume — tap.delay~/tap.multitap~ are the honest tools.
  • You want more than five strings. The count is fixed; mc. wrapping the whole bank gives you choirs of banks, at which point you are building a sympathetic-string instrument and should budget CPU like it.

The part that comes apart, on tape

Three objects and one posture: hands on the controls while it runs. The stutter takes a phrase apart, the tape echo smears the pieces, and the fuzz decides how hard the whole thing is being pushed. None of the three has a "right" setting, which is the point of the part of the book they live in.

This recipe is a rig, not a record. What is documented about the Kid A-era working method is that a laptop running Max sat in the signal path and got played — the objects here are informed by what that rig was for, not by anyone's patch. Nothing below is claimed to be a reconstruction of a specific track.

Measured claims are borrowed from four heads and a motor, the part that comes apart and the dirt with two stages.

The chain

source ──▶ tap.fuzz~ ──▶ tap.stammer~ ──▶ tap.tapecho~ ──▶ out

Order matters, and this order is the useful one:

  • Fuzz first, because a stutter of a distorted signal is a stutter; a distortion of a stuttered signal turns every slice edge into a transient the clipper amplifies.
  • Echo last, because the echo is the only object here that is supposed to blur. Put it before the stutter and the stutter slices the blur, which sounds like a mistake rather than a decision.

The dirt

Keep it low. Two saturators in series get muddy fast, and the echo has its own drive.

controlsetting
gain0.35
edge0.3
asymmetry0.2
bass / treble / contrast0. / 0.1 / 0.3
oversample4
levelto taste, usually negative

contrast is the scoop, and a scooped source stutters better than a mid-heavy one — the slices stop fighting the vocal or the guitar they came from. Leave oversample at 4 and resist the urge to save the cycles at 2: a hard edge on bright material is exactly the case where one doubling stops being enough, and 2 measures badly above about 6 kHz.

The stutter

controlsettingwhat it does
step60. msthe grid
divisions1
density0.3how often it grabs
repeats4how long it holds
reverse0.2chance a repeat plays backwards
jump250. mshow far back a slice may reach
fade2. msthe flanks
seedany integer
mix100

The one thing to internalize: repeats is the hold, density is the grab. Occupancy — how much of the timeline has a slice in flight — measures 41 % at density 0.3 / repeats 1 and 90 % at density 0.9 / repeats 1, but density 0.3 with repeats 6 already sits at 76 %. If the part feels too busy, pull repeats before you touch density; you will keep the sparseness of the entrances while shortening what each one does.

seed is a contract, not a flavour: the same seed and the same moves give a bit-identical render, and two instances on two tracks decorrelate by seed alone. Different seeds change 89 % of samples, so it is a real dice roll, not a phase tweak.

At density 0. the object is a bitwise bypass — worth knowing, because it means you can automate density to zero and get the dry signal back exactly, with no crossfade artifact to work around.

The material contract. This object flatters a played phrase and flatters a sustained note far too much. Slice similarity measures 1.000 on a held sine against 0.286 on a plucked phrase: on sustained material every slice is interchangeable, so the stutter has nothing to expose and sounds like a tremolo. Feed it something with transients and pitch variety.

The tape

controlsetting
span400. ms
heads3
ratios0.25 0.5 1.
levels0.7 0.85 1.
pans0.2 0.8 0.5
regen0.55 (the ride starts here)
darken5000.
drive0.4
wow / flutter0.4 0.35 / 0.3 8.
mix35

span is defined at the ratio-1.0 head, so the head at 1. returns at 400 ms and the others at 100 and 200. Change span while it runs and the whole thing glides like tape speed rather than splicing — that is the transport, and it is the second-best gesture in this rig.

The best one is regen. It goes past 1 on purpose. Past unity the line self-oscillates, and it stays bounded because the saturator caps what comes back: the ceiling is |in|max + regen/drive, measured under that value at every drive tried. So a ride up to 1.4 and back is a controlled build, not a fire. Keep drive up while you do it — drive 0. removes the saturator and the cap falls back to unity, which is the setting where a long ride will not behave.

darken is the generation loss, and it is what makes repeats decay into a shape instead of just getting quieter. 5 kHz is a good default; below 3 kHz the tail turns to mud, which is sometimes what you want under a chorus.

The gestures, in order of value

  1. Ride regen past unity and back, while mix stays put. One hand, whole arrangement.
  2. Automate density to 0 and back. Exact bypass, so it reads as the part reassembling rather than a fade.
  3. Move span during a held note. Varispeed glide, not a splice.
  4. Change seed between takes, never during one.
  5. reverse and jump. Character, and cheap to overdo. Note jump is milliseconds, not a probability — at 250 the machine starts quoting material from a quarter-second before the slice it just took, which is where a stutter stops sounding like a stutter and starts sounding like an edit.

When to leave the recipe

  • The source is sustained. See the material contract. Put the stutter on the drums and leave the pad alone.
  • You want the slices in time with something. Nothing here syncs to a transport; step is milliseconds. Drive it from your own clock if you need bars.
  • You want the echo to stay clean. drive 0. gets you a clean line — but then do not ride regen past 1, because the cap that makes that safe is the saturator you just removed.
  • You want the dirt to be the point. Then the fuzz belongs last, not first, and this is a different recipe: tap.tapecho~ → tap.fuzz~ with the echo's own drive at 0. Distorting a wash is a real sound; it is just not this one.

Two hands on a live buffer

tap.scrub~ is the object in this library that most needs a controller attached before it means anything. Everything below assumes one: an XY pad, two faders, a trackpad, a phone sending OSC — anything that gives you two continuous values at once. The recipe is mostly about what to put on each axis and why.

Measured claims are borrowed from two hands on the same tape.

The chain

source ──┬─▶ tap.scrub~ ──▶ tap.palme~ ──▶ out
         └────── (its own dry path, via mix) ──────▶

The scrub records what passes through it, so it goes in the signal path rather than on a send — there is nothing to send it that it is not already hearing. Its mix is the dry/wet, and at mix 0 it is the input bit for bit, which means you can leave it in the chain permanently and have it be audibly absent until you touch it.

The pad

controlsetting
maxhistory4. s (object argument; bought at DSP start)
positionX axis, 0–1500 ms
pitchY axis, −12 to +12 st
drift0.
size80. ms
overlap2
spray0.
mix100 while playing
smooth0. if driving by signal, 20. if by messages

X is position, Y is pitch, and the whole object exists because those are independent. On tape they would be the same axis — moving the head is the pitch change. Here you can rake back through the last second and a half at the pitch you started at, or hold still and transpose, or do both at once in different directions. Spend the first five minutes doing each separately; the object does not become obvious until you have felt that they do not interact.

size is the texture control. 80 ms is a granular pad. Down at 20 ms it turns metallic and starts pitching itself at the grain rate; up at 200 ms it stops being granular and becomes a soft varispeed. Sizes that divide evenly by overlap have an exactly flat window sum — 80 with overlap 2 does — which matters when you want the still position to be clean.

overlap 1 is a texture, not a mistake. It leaves gaps between grains: a gated, chopped version of the same gesture. Worth a switch on the controller.

The honest bit about pitch

Transposing here warbles, and it is measured rather than apologized for: 98.8 % of a perfect shifter's energy lands within ±15 Hz of the transposed pitch (worst case 91.7 %), so the note is right — what the object loses is concentration, 92.0 % as focused as a clean shift and 75.0 % at worst. Audibly that is a warble, and it is the classic single-delay-line pitch-shifting artifact rather than anything peculiar to this kernel.

Two ways to work with it:

  • Lean in. spray 30. trades the narrow comb for a broadband smear. On sustained material this reads as a texture rather than a fault, and it is the better answer for pads.
  • Stay out of its way. Keep the Y axis to ±7 and let the position do the work. Small intervals warble least and the object is a scrub pad first.

If you need a clean shift, this is the wrong object — tap.shift~ is built for it. (tap.pitchaccum~ is built for shimmer rather than transparency, and has an open issue about where its line actually sits.)

Freeze, which is the other half

freeze 1 stops the recorder. The playhead keeps going, so the position now addresses fixed tape and the grains loop the same window — and you can still scrub, transpose, drift and spray through it. Nothing is going into the input any more, which is the point: it is a hold you can perform.

A sequence that works on stage:

  1. Play the phrase through at mix 0. Nothing happens; the tape fills.
  2. mix 100, freeze 1. The last few seconds are now the instrument.
  3. Drag X slowly. This is the scrub.
  4. drift -0.3. The playhead walks backwards on its own while you keep your hand free for Y.
  5. spray 40., size 200.. It stops being a phrase and becomes a pad.
  6. freeze 0, mix 0. The room comes back.

Step 6 is exact — mix 0 is bitwise passthrough — so the return is clean however far out step 5 went.

The one constraint

A grain born position behind the live edge and playing at rate r reaches position − size·(r−1) behind it by its end. Transpose up with the position near the live edge and the grain's tail runs off the front of the tape into the oldest material — a seam.

Nothing clamps it, because clamping would silently bend the pitch to keep the grain in bounds, which is a worse failure than the seam. Practically: keep the position at least size·(rate−1) back. At size 80 and one octave up that is 80 ms. Setting the X axis to start at 100 ms rather than 0 makes the whole problem disappear, and this is why the table above says 0–1500 rather than 0–1500 starting at zero.

What each ingredient buys, in order

  1. A controller with two continuous axes. Without it this is a delay.
  2. freeze. The half of the object you can build a performance on.
  3. size. The texture, and the only control that changes what kind of thing you are playing.
  4. A diffuseur after it. tap.palme~ @mix 40 sustains what the scrub chops; the strings fill the gaps that overlap 1 opens.
  5. spray. Trades one artifact for another. Real, and last.

When to leave the recipe

  • You want it in time. Nothing here syncs. That is tap.stammer~, on the same tape — literally the same capture code — and the two are meant to be swapped between rather than combined.
  • You want a clean delay. tap.delay~ costs a fraction as much and windows nothing.
  • You want the position to feel like a jog wheel. Put drift on a spring-loaded control and leave position alone: drift is velocity where position is location, and for wheel-like gestures velocity is the right variable.

The instrument in the corner of the room

The Ondes Martenot is not a synthesizer, and the fastest way to make it sound like one is to patch it like one. This recipe assembles the four objects that make up the actual instrument — voice, key, and a loudspeaker with a body — and then spends most of its length on the part that is not a setting at all: what your two hands do.

Everything measured here is borrowed from the instrument that is not a synthesizer and loudspeakers you can play, which in turn cite the circuit paper (Najnudel, Hélie, Roze & Boutin, IEEE/ACM TASLP 28, 2020) and the intensity-key measurement (Quartier et al., Acta Acustica 101(2), 2015).

The chain

[ribbon signal] ──▶ tap.ondes~ ──▶ tap.palme~ ──▶ out
[key signal]    ──▶     ▲              (or tap.metallique~)

Two signals in, one instrument out. That is the whole rig, and the temptation to put things between the stages should be resisted until you have played it as it stands — the voice and the diffuseur were designed to be adjacent, and every stage you insert is a stage the real instrument does not have.

The voice

controlsettingwhy
ribbondriven by signalsemitones above A1, not Hz
keydriven by signal0–1 of the physical travel
depth1.equal oscillators; the full harmonic series
detect0.2the published R4 × C21, 200 µs
drive1.nominal; the harmonics are already there
keyplacement0pressure is level
polarity1
power0the 2A3 moves total harmonics by 0.003
oversample4
smooth0.the signal inlets are not ramped anyway

Start there and change exactly one thing at a time, because most of these are citations rather than tastes and the object will tell you when you have left the instrument behind.

The two that are genuinely yours: keyplacement 1 moves the key in front of the valves, so hard presses get dirty as well as loud — worth about 0.09 of total harmonic content at a half-press, and it is the single change that most makes the object feel like a synthesizer rather than an ondes. polarity -1 flips which side of the waveform the preamplifier bends, worth about 0.12. Try both; keep whichever suits the piece.

depth below 1 is the cheapest real timbre move in the object. At 0.4 the envelope never closes and the tone thins toward a sinusoid — the closest thing here to a "register", and it is a physical mismatch between two oscillators rather than an invented control.

The hands

This is the recipe.

The ribbon is linear in semitones, because the circuit paper's Eq. 7 makes it so. That single fact is why an ondes glissando sounds like an ondes glissando: a hand moving at constant speed produces a constant-rate glide, not the accelerating swoop a linear-in-Hz control gives you. So drive ribbon with something that moves linearly in semitones over time:

line~ 0. 36. 4000   ──▶ tap.ondes~ left inlet

Three octaves in four seconds, and it will sound even the whole way. Build your phrases the same way — line~ or a slow sig~ ramp per note, never a quantized step unless you specifically want the instrument to sound wrong. Nothing here rounds to a semitone, and that is deliberate.

The key is the dynamics, and it starts silent. Roughly the bottom 45 % of the travel makes no sound at all. That is the key's own first phase, the elastic strip bending before it reaches the powder bag, and it is why the instrument attacks so sharply: the whole 50 dB lives in the 4.5 mm right after the silence. Practically:

  • Drive key from a pedal, a fader, or a line~ — anything continuous.
  • Expect nothing below about 0.45. If you want the note to speak the instant your controller leaves zero, rescale: scale 0. 1. 0.45 1.. Do that only if you want to give up the attack, because that dead travel is what lets you place an entrance to the millisecond.
  • The curve steepens through the middle and flattens at the top. Crescendos therefore want a decelerating controller move, not a linear one — which is exactly the feedback a player's finger gets from the real spring.

Play them together. The ribbon without the key is a test tone; the key without the ribbon is a volume pedal. The instrument is the two hands, and ondes_ribbon.wav in the render set exists to demonstrate what that sounds like when both are moving.

The loudspeaker, which is an instrument too

tap.palme~ is the default answer. Twelve strings on a board, and they sustain what the voice has already stopped playing:

controlsetting
root110. — put the lowest string under your part's key
tuning0 chromatic, so the board answers every note
decay6.
damping4000.
detune4. — cents of scatter, so the board is not a chorus unit
drive / asymmetry / saturation1. / 0.3 / 0.2
mix60

tuning 1 puts the harmonic series on the root instead: the board then answers only what belongs to that key, which turns a chromatic line into something that blooms on some notes and stays dry on others. Use it when the piece really is in one key, and hear it as a compositional decision rather than a preset.

Watch the level. Twelve resonant loops add up, and a driven board can be a great deal louder than what went into it — level is there for that, and it is the one control on these objects you will need to touch first.

tap.metallique~ is the other cabinet: eight plate modes rather than twelve strings, so it colours instead of harmonizing. @pitch 180 @decay 6 @tilt 0.8 @brightness 1. @mix 50 is the gong. Push drive to 3 with asymmetry 0.5 and the distortion happens before the plate, because that is where the transducer is — a distorted waveform ringing a gong, not a distorted gong. It is not subtle and it is the most distinctive sound in the family.

What each ingredient buys, in order

  1. tap.ondes~ with both hands moving. Everything else is optional. A static ribbon and a static key is a demo, not an instrument.
  2. The diffuseur. The voice alone is thin on purpose — it is a valve preamplifier output, and it was never meant to be heard without a body after it.
  3. depth. The one timbre control that costs nothing and is physical.
  4. keyplacement / polarity. Real, measured, and yours to choose.
  5. power. Measured at 0.003 of total harmonic content. Last.

When to leave the recipe

  • You want the waveform registers. The real instrument has switchable timbres — creux, gambe, nasillard and the rest. They are not here, and they are not here on purpose: no source obtained describes their filter shapes, and inventing them is the one thing these objects will not do. If you need those colours, put a filter after tap.ondes~ and call it your filter, not Martenot's.
  • You want polyphony. The instrument is monophonic and so is this. Two tap.ondes~ in parallel is a duet, not a chord — which is how ondes ensembles actually worked, so it is not a bad answer.
  • You want the diffuseurs on something else. Take them. A guitar into tap.palme~ is the best argument for shipping them standalone, and nothing about them needs the voice in front.
  • You want tap.triode~ on its own. It is a published valve stage and it works on anything: @tube 2 @stage 2 @drive 6 is the 2A3 power stage used for something it was never in this instrument for.