How I reverse engineered a commercial spatial audio effect

Hacker News by 8 min read 47x views
How I reverse engineered a commercial spatial audio effect

Share Post

Back during I was using a distinct OS, I used a program which processed the audio being played in genuine period and made it audio much better. It made the audio awareness farther distant and a bit lighter to hear, during carrying the applicable information. I’d continually wanted a akin program which was likewise uncomplicated to use for Linux, and decided to compose one.

0:00 / 0:10

Listen alongside headphones. The excerpt is from P.I.M.P. by Bacao Rhythm & Steel Band[0].

The processed type appears farther from the ears and feels small “on the ear” making it awareness small “blocked” following listening to music for a while.

First attempt

Since I was penning for Linux, my basis stack was PipeWire[1]. I had no previous audio experience, so I began by using an LLM as a sounding commission during experimenting alongside norm audio filters and psychoacoustic techniques.

Somewhere alongside the way I learnt concerning HRTFs, which depict how audio arriving from a particular direction reaches the ears, depending on the direction and geometry of the listener. I tried the MIT KEMAR database[2] archetypal and idea it echoed awful, since it had been recorded in an anechoic chamber, so there were no area echoes, which were a desirable psychoacoustic consequence here. The outcome echoed cheap and additional inwards to the ear. Later I established the SADIE II database[3], which additionally came alongside area responses, and that echoed much nearer to what I remembered.

But this method shortly showed its limits. The LLM began talking concerning things akin “perceptual envelopment”, “late-field decorrelation”, “spectral diffusion” and “diffuse spatial energy”.[1][1] Gemini was particularly fine at this. It was additionally the lone one alongside a liberated pupil program. While I could comprehend that item was wrong, those descriptions didn’t item to item I could comprehend or use to enhance the effect.

I remaining the project deceased for a during alongside no apparent direction to continue with.

Re-attempt

Much later, I had a distinct idea. Instead of trying to detect the exact parameters or techniques and algorithms the program used internally, I could study the behavior of the display itself.

This seems apparent in retrospect, but I did not have a copy of the program. I could lone test on a friend’s laptop, so I could not keep changing my implementation and comparing it alongside the citation whenever I wanted.

I began study Signalsmith[4] and several blogs by audio engineers to build adequate intuition to cognize what I should measure, and shortly realised that I didn’t need to keep using the program to grasp its encourage responses and another measurable features. I generated an audio document on Linux, got my friend’s laptop for a while, played the generated document through the program and recorded the output.

The gathering of tests

The inquiry became what I should put into that document so it could grasp the applicable particulars which differentiated the program without having to re-record features or filters.

The archetypal helpful item I established was impulse-response measurement.

If I continue a document containing a sole nonzero example through the filter, everything which appears following it in the record would’ve been added by the filter. Features specified as the degree of its decay complete time, the form of the decay and the phase delays connected alongside it would be current in the output. Playing the encourage in the correct conduit allows extracting the two outputs’ responses to audio played on the right, and likewise for the left.

If the reply to one encourage is $h$, afterward the reply to an encourage at stance $k$ is the identical $h$ shifted to $k$ and scaled by $x[k]$. Adding all of those responses gives

$$ y[n]=\sum_k x[k]h[n-k]=(x*h)[n]. $$

This of way depends on the display being a linear time-invariant filter[5].

A louder encourage should create the identical reply made proportionally louder:

$$ F(ax)=aF(x). $$

Two inputs played together should create the sum of what all input produces separately:

$$ F(x_1+x_2)=F(x_1)+F(x_2). $$

And playing the identical input afterward should create the identical reply later.

And stated it’s stereo, the complete equation would be:

$$ \begin{bmatrix} y_L\\ y_R \end{bmatrix} {}={} \begin{bmatrix} h_{LL} & h_{LR}\\ h_{RL} & h_{RR} \end{bmatrix} * \begin{bmatrix} x_L\\ x_R \end{bmatrix}. $$

Given I was sending one document to have it recorded through all mode, I added impulses at distinct amplitudes to inspect for level-dependent processing, and played the identical encourage in the two channels, afterward alongside the correct conduit inverted. The outputs should equivalent the secluded responses added together in the archetypal case and subtracted in the second, checking the conditions above. I additionally remaining adequate silence between them for the reply to complete before the next impulse.

Impulses at three levels, played on the left, right, the two channels and alongside contrary signs.

While study concerning this, I noticed that document on measure tended to use sine sweeps fairly than impulses[6].[2][2] An encourage already contained all frequency, but having lone one example meant there was small energy at all frequency, so a lengthy clear would create the reply easier to differentiate from record noise.

A shortened logarithmic sweep.

$$ Y(f)=H(f)X(f), \qquad H(f)=\frac{Y(f)}{X(f)}. $$

This would activity for most of the range, but if $X(f)$ was near to zero, dividing record noise by it would create fairly a notable display which had small to do alongside the program. This could be avoided alongside a regularized version.

$$ \hat H(f)=\frac{Y(f)X^*(f)}{|X(f)|^2+\lambda}. $$

With adequate input energy this would about be the identical division, during $\lambda$ would bounds the outcome elsewhere. The complex values additionally keep the phase, which was required stated audio sent to one conduit could appear in the another alongside a delay.

The clear gave another measure to difference alongside the impulse, but it additionally took much longer to play, during which the program could alter its behaviour. I added dependable tones at distinct frequencies and amplitudes, and longer tones going from quiescent to noisy and back, to see whether the acquire changed and how lengthy it took to do so.

The dependable tones, shortened and played through the two channels.
A 1000 Hz spirit alternating between two levels.

And a multitone to inspect if any new frequence pops up:

The eight-frequency multitone.

And repeat to inspect for time-invariance:

Three identical copies of the multitone.

And I kept a division of noise aside, to test whether the extracted display could reproduce its recorded output afterwards.

Noise covering the low, center and elevated frequence ranges in sequence.

Recovering the filter

Given record and playback were started manually, the record archetypal needed an offset. I had put evenly spaced clicks near the start, which I initially tried finding by correlating the record alongside the first clicks. Since the display could dispersed all click complete time, I summed the squared samples complete a opening $W$ to evaluation the energy about it instead,

The eight synchronization clicks.

$$ e[n]=\sum_{m\in W}\left(y_L[n+m]^2+y_R[n+m]^2\right), $$

and established the offset through

$$ \hat\tau=\arg\max_\tau\sum_n e_{\text{probe}}[n]e_{\text{recording}}[n+\tau]. $$

I applied the identical offset to the two channels so the postpone between them was retained.

Output

For the chief encircle setting, the scaled impulses were nearly identical, the blended inputs matched the sum of their distinct responses, and the multitone produced nearly no new frequencies. To difference the reconstructed output alongside the recording, I used

$$ \varepsilon=\frac{\lVert y-\hat y\rVert_2}{\lVert y\rVert_2}, $$

where $y$ was the recorded output and $\hat y$ the output produced by the extracted filter. Given the identical repeated sections already differed slightly, I compared the model’s error alongside that difference,

$$ \varepsilon_{\text{model}}\approx\varepsilon_{\text{repeat}}\ll1. $$

The volume-reducing manner changed its acquire alongside amplitude and took period to reconstruct it, during another spatial manner changed between repeats, so neither could be reproduced alongside a fixed convolution filter.

Recording the two output channels for an encourage in all input conduit gave me the four filters.

Sound additionally appeared in the contrary channel, quieter and slightly delayed, followed by a lengthy tail. Together these contained the out-of-head consequence I had been trying to create earlier.

To decide how much of this reply to keep, I tried increasingly lengthy filters against the noise section,

$$ \varepsilon_N= \frac{\lVert y_{\text{noise}}-h^{(N)}*x_{\text{noise}}\rVert_2} {\lVert y_{\text{noise}}\rVert_2}, $$

and kept the complete response, which predicted the noise additional accurately than the shorter versions. The reply extracted from the clear additionally accepted complete the range anywhere it had adequate input energy.

Loading the four responses into Saq[7] gave me the placement I had been trying to reproduce by hand.

You can try the outcome here. Choose your own local audio document or grasp a tab (when not in a Firefox browser[8]), afterward toggle encircle to compare.

Bibliography

Other Article Hacker News
↑
Close Right Ads
Close Left Ads