Feedback suppression

Vicente González Ruiz & Savíns Puertas Martín

October 8, 2026

Contents

1 Roadmap
2 The problem
3 The trivial (and most efective) solution
4 A trivial (but unpractical) solution
5 The “simplest” (but in general, ineffective) solution
6 Improving the performance through filtering
6.1 Filtering in the frequency domain
7 Taking turns speaking?
8 Deliverables
9 Resources

1 Roadmap

In this milestone, we’ll solve the problem generated by the recording of the signal emitted by our speakers that is captured by our microphone, generating an unpleasant feedback that difficults our conversation (using InterCom). First, we’ll formalize the problem and then we will explore different solutions, with varying effectiveness and computational requirements.

2 The problem

One of the first problems we encounter with the use of the buffer.py module1 is that, if we don’t use headphones, the sound that comes out of our PC’s (loud)speaker some time (miliseconds) later reaches our mic(rophone), and more some time later, that sound reaches our interlocutor (the “far-end” ... in the system we are the “near-end”) in the form of an echo (signal) of its own voice, which is reproduced by his/her speaker, which can be captured again (some time later) by his/her mic and sent it back to us ... and so on, generating a rather unpleasant feedback “noise”.

To formalize this problem, let’s define:

  1. \(s\) as the (analog) signal played by our speaker,
  2. \(n\) the signal emited by the near-end person (that’s me)2, and
  3. \(m\) the (mixed) signal recorded by my microphone3 ,

where

\begin{equation} m(t) = \tilde {n}(t) + \tilde {s}(t), \label {eq:echo_problem} \end{equation}

where \(\tilde {\cdot }\) represents an approximation of \(\cdot \). We have here approximations because the signals are modified when they travel through the air.

Our problem here is to minimize the energy of \(\tilde {s}\), i.e., to make

\begin{equation} \tilde {s}(t) = 0. \end{equation}

3 The trivial (and most efective) solution

Use a headset. In this case,

\begin{equation} m \approx n \label {eq:headset_solution} \end{equation}

because in this case, \(s(t)\approx \tilde {s}(t)\approx 0\).

4 A trivial (but unpractical) solution

Decrease the gain of the amplifier of your speaker to do (the energy of) \(\tilde {s}\) as small as possible. Unfortunately, this also decreases the volume of the far-end signal \(s\) (the voice of our interlocutor) :-/

5 The “simplest” (but in general, ineffective) solution

Lets \(\mathbf m\) the digital version of \(m\), and \({\mathbf m}[t]\) it’s \(t\)-th sample4. In this “simplest” solution, we send

\begin{equation} \tilde {\mathbf n}[t] = {\mathbf m}[t] - a{\mathbf s}[t-d], \label {eq:simplest} \end{equation}

where \(a\) is an attenuation (scalar) value, and \(d\) represents the delay5 (measured in sample-times) required to propagate the sound waves from our speaker to our mic. We define

\begin{equation} \hat {\mathbf s}[t] \triangleq a{\mathbf s}[t-d] \label {eq:minimal_filter} \end{equation}

as the estimated6 feedback signal that reaches our microphone at the same instant of time that the sample \({\mathbf n}[t]\) would have been captured in the ausence of the feedback.

Notice that it have been used the notation \(\hat {\cdot }\) to emphasize that \(\hat {\mathbf s}\) is a (“registered”7) prediction for \(s\) reaching our microphone. Notice also that, if \({\mathbf s}[0]\) is the first sample of a chunk (for example, the \(c\)-th chunk), the sample \({\mathbf s}[-d]\) could belong to a previous chunk (the \((c-1)\)-th chunk).

Finally, \(a\) should be choosen considering that under ausence of voice in each end, \(s(t)\approx 0\). For example, Skype estimates \(d\) and \(a\) using a “call-signal” (a sequence of more-or-less tonal sounds). \(d\) is determined measuring the propagation time of the call-signal between our speaker and our mic.

Why it’s difficult to make it work?

This algorithm is ineffective because:

  1. \(d\) can change over time (for example, if we’re using a laptop and move the screen, \(d\) will vary). This can be costly, because even using correlation in the Fourier domain, this is a heavy operation to be performed in real-time.
  2. Assuming that the signal captured by the microphone is \(a{\mathbf s}[t-d]\) is an oversimplification. This signal is only similar (usually it’s filtered by the environment and likely contains echoes).

6 Improving the performance through filtering

We can improve the performance of the previous feedback cancellation solution if we take also into consideration that the feedback signal \(\tilde {s}\) that finally reaches our microphone is (at least in part) the convolution of \(s\) and a signal \(h\) that represents the echo response of our local audioset (speaker, mic, walls, monitor, keyboard, our body, ...) to an impulse signal \(\delta (t)\).8 In other words, we can modify Eq. \eqref{eq:simplest} to compute

\begin{equation} \tilde {\mathbf {n}}[t] = \mathbf {m}[t] - (\mathbf {h}\ast \mathbf {s}[t-d]) = \mathbf {m}[t] - (\mathbf {h}\ast \mathbf {s})[t-d], \label {eq:using_convolution} \end{equation}

where \(\ast \) represents the (digital) convolution between (in our case of) digital signals, and \(\mathbf h\) is the digitalized version of \(h(t)\).

6.1 Filtering in the frequency domain

The convolution of digital signals in the time domain can be expensive (with computational complexity \(O^2\), where \(O\) is the number of elements to process) if the number of samples or/and filter coefficientsis is high. Fortunately, thanks to the convolution theorem [1, 2], the convolution can be replaced by the dot product (with complexity \(O\)), when we consider the signals in the frequency domain. Thanks to this, we can rewrite the Eq. \eqref{eq:using_convolution} as

\begin{equation} \tilde {\mathbf n}[t] = {\mathbf m}[t] - ({\mathcal F}^{-1}\{{\mathbf S}{\mathbf H}\})[t-d], \label {eq:faster} \end{equation}

where \(\mathbf S\) is the (digital) Fourier transform9 of \(\mathbf s\), \(\mathbf H\) is the Fourier transform10 of \(\mathbf h\), and \({\mathcal F}^{-1}\) represents the inverse (digital) Fourier transform. Notice that all these transforms are applied to digital signals, and there exist fast algorithms (with complexity \(O\log _2O\)) to “travel” from the signal domain to the frecuency domain, and viceversa.

Practical issues

Unfortunately, even when we expect that this improved feedback supression algorithm is going to perform better than the previous one, the computation of the filter weights \(\mathbf h\) requires emitting impulses that can be heard by the user. Furthermore, while the acoustic response of the environment is being analyzed, the near-end should remain silent; otherwise, its own voice would be treated as an echo and attenuated by the filter.

7 Taking turns speaking?

Lets recap. Our problem is that, without a headset, we have that the signal that we send is

\begin{equation} \mathbf {m} = \tilde {\mathbf {n}} + \hat {\mathbf {s}} \end{equation}

where \(\tilde {\mathbf {n}}\) is the version of our voice captured by our mic, and \(\hat {\mathbf {s}}\) is a aproximated-and-multy-echo version of \(\mathbf {s}\), the signal played by our speaker(s). And if we were able to make \(\hat {\mathbf {s}}=\mathbf {0}\), our feedback problem would vanish, ... at least, theoretically.

But wait ..., as a polite person, if I don’t speak when I am listening to (my interlocutor), I could assume that \(\hat {\mathbf {s}}=\mathbf {0}\) when I am speaking because my interlocutor should do the same (be silent when I speak)! But, if this is not true? A solution is: if I am speaking, can attenuate \(\mathbf {s}\) (the signal played by my speaker).

Summarizing: you should be able to control the volume of \(s\) (the analog signal that comes out from my speaker) and cut it down (or even off, if necessary) when you are speaking or your mic is recording a signal \(m\) with enough energy.

8 Deliverables

  1. Write a Python module, called feedback_supression.py, that inherits from buffer.py and that implements one of the proposed solutions (or any other you prefer). Define a new command-line parameter to switch on this functionality, that by default should be disabled (off). If necessary, use extra parameters that can be controlled in real time.
  2. Insert the code of feedback_supression.py in a Jupyter notebook cell and generate the archive feedback_supression.py using %%writefile feedback_supression.py. The notebook should also include the names of the members of the working group and I example of how to use your implementation.
  3. Notice that for testing your implementation it is essential that your operating system does not cancel the echo. If your OS cancel the echo, you must perform the tests in Linux (see Framework).

9 Resources

[1]
J. Kovačević, V.K. Goyal, and M. Vetterli. Fourier and Wavelet Signal Processing. http://www.fourierandwavelets.org/, 2013.
[2]
Alan V. Oppenheim, Alan S. Willsky, and S. Hamid Nawab. Signals and Systems (2nd edition). Prentice Hall, 1997.