Three separate costs, found with a sampling profile of the RX chain with
emnr forced on:
aepf() averaged mask[] over a window of N = 2*psi + 1 = 41 bins by walking
the window for every one of the 2049 bins, i.e. O(msize*N). Its three spans
are all symmetric windows clipped at the array ends, so take each from a
prefix sum instead: one subtraction per output, O(msize).
xemnr() advanced four ring indices with a '% size' per step. iasize is 4096
and oasize 1024 here, and neither is known to the compiler, so each step was
a real integer division -- ~8700 of them per frame. The indices step by one
and, since iasize >= fsize and oasize >= incr always hold, wrap at most once
per loop, so walk contiguous runs and wrap between them.
calc_gain() called getKey() twice per bin with the same gamma, so the gamma
row index and its log10 were computed twice. Split getKey into keyIndex() +
keyLerp() and locate gamma once. The remaining logs go through wdsp_log10()
(new fastmath.h), accurate to 2e-13 against libm and ~2.4x its throughput;
gamma and xi are bracketed against the table limits first, so the argument
is always positive and normal. Also clamp the row index so the second
bilinear corner cannot address the next row of the 241x241 table.
Measured in situ on an Apple M1 Pro, 512-sample buffers, cost of turning
emnr on, best of 5:
baseline 73661 ns
+ aepf, ring walks 48927 ns 1.51x
+ getKey 37852 ns 1.95x
Output is not bit-identical, as the prefix sum and the reassociated logs
round differently. Over 300 buffers with emnr alone the worst deviation is
4.0e-09, an SNR of 196 dB; perturbing a single input sample of the unmodified
code by one ulp diverges it from itself by 1.4e-08 (186 dB), so this change
disturbs the chain less than the last bit of the input does.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The tap loop wrapped the ring with a test on every tap:
if ((idx_out = idx_in + j) >= ringsize) idx_out -= ringsize;
which made the address non-affine and stopped the vectorizer. Split the
walk at the wrap point instead, so both halves are unit-stride, and split
the complex ring into separate I/Q arrays so the taps load contiguously
rather than through a de-interleaving ld2.
The dot product also could not be vectorized as written: reassociating an
fp reduction needs -ffast-math, which this library must not enable (it
relies on IEEE semantics for 0/0 = NaN and x/0 = Inf, see linux_port.h).
Carry four independent accumulator pairs instead, which both breaks the
FMA dependency chain and lets the vectorizer in on any compiler.
Measured on an Apple M1 Pro, 512-sample DSP buffers, best of 5:
xresample, decimation to 48 kHz before after speedup
192k -> 48k (561 taps) 342.0 us 84.2 us 4.06x
384k -> 48k (1121 taps) 708.9 us 175.1 us 4.05x
576k -> 48k (1681 taps) 1074.4 us 266.9 us 4.03x
768k -> 48k (2241 taps) 1438.9 us 360.0 us 4.00x
full xrxa() chain, 576k input 1148.2 us 337.6 us 3.40x
10.76% 3.16% of one core
xresampleF (float, host audio) 2.6x - 3.4x
Summation order changes, so the double path is not bit-identical: over
400 buffers of the full RX chain the worst deviation is 1.1e-12, an SNR
of 251 dB. The float path is bit-identical, as the cast to float absorbs
the difference. With the resampler bypassed (48k in, 48k out) the chain
is unchanged bit-for-bit.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>