In this milestone, we’ll solve the problem generated by the recording of the signal emitted by our speakers that is captured by our microphone, generating an unpleasant feedback that difficults our conversation (using InterCom). First, we’ll formalize the problem and then we will explore different solutions, with varying effectiveness and computational requirements.
One of the first problems we encounter with the use of the buffer.py module1 is that, if we don’t use headphones, the sound that comes out of our PC’s (loud)speaker some time (miliseconds) later reaches our mic(rophone), and more some time later, that sound reaches our interlocutor (the “far-end” ... in the system we are the “near-end”) in the form of an echo (signal) of its own voice, which is reproduced by his/her speaker, which can be captured again (some time later) by his/her mic and sent it back to us ... and so on, generating a rather unpleasant feedback “noise”.
To formalize this problem, let’s define:
where
where \(\tilde {\cdot }\) represents an approximation of \(\cdot \). We have here approximations because the signals are modified when they travel through the air.
Our problem here is to minimize the energy of \(\tilde {s}\), i.e., to make
Use a headset. In this case,
because in this case, \(s(t)\approx \tilde {s}(t)\approx 0\).
Decrease the gain of the amplifier of your speaker to do (the energy of) \(\tilde {s}\) as small as possible. Unfortunately, this also decreases the volume of the far-end signal \(s\) (the voice of our interlocutor) :-/
Lets \(\mathbf m\) the digital version of \(m\), and \({\mathbf m}[t]\) it’s \(t\)-th sample4. In this “simplest” solution, we send
where \(a\) is an attenuation (scalar) value, and \(d\) represents the delay5 (measured in sample-times) required to propagate the sound waves from our speaker to our mic. We define
as the estimated6 feedback signal that reaches our microphone at the same instant of time that the sample \({\mathbf n}[t]\) would have been captured in the ausence of the feedback.
Notice that it have been used the notation \(\hat {\cdot }\) to emphasize that \(\hat {\mathbf s}\) is a (“registered”7) prediction for \(s\) reaching our microphone. Notice also that, if \({\mathbf s}[0]\) is the first sample of a chunk (for example, the \(c\)-th chunk), the sample \({\mathbf s}[-d]\) could belong to a previous chunk (the \((c-1)\)-th chunk).
Finally, \(a\) should be choosen considering that under ausence of voice in each end, \(s(t)\approx 0\). For example, Skype estimates \(d\) and \(a\) using a “call-signal” (a sequence of more-or-less tonal sounds). \(d\) is determined measuring the propagation time of the call-signal between our speaker and our mic.
This algorithm is ineffective because:
We can improve the performance of the previous feedback cancellation solution if we take also into consideration that the feedback signal \(\tilde {s}\) that finally reaches our microphone is (at least in part) the convolution of \(s\) and a signal \(h\) that represents the echo response of our local audioset (speaker, mic, walls, monitor, keyboard, our body, ...) to an impulse signal \(\delta (t)\).8 In other words, we can modify Eq. \eqref{eq:simplest} to compute
where \(\ast \) represents the (digital) convolution between (in our case of) digital signals, and \(\mathbf h\) is the digitalized version of \(h(t)\).
The convolution of digital signals in the time domain can be expensive (with computational complexity \(O^2\), where \(O\) is the number of elements to process) if the number of samples or/and filter coefficientsis is high. Fortunately, thanks to the convolution theorem [1, 2], the convolution can be replaced by the dot product (with complexity \(O\)), when we consider the signals in the frequency domain. Thanks to this, we can rewrite the Eq. \eqref{eq:using_convolution} as
where \(\mathbf S\) is the (digital) Fourier transform9 of \(\mathbf s\), \(\mathbf H\) is the Fourier transform10 of \(\mathbf h\), and \({\mathcal F}^{-1}\) represents the inverse (digital) Fourier transform. Notice that all these transforms are applied to digital signals, and there exist fast algorithms (with complexity \(O\log _2O\)) to “travel” from the signal domain to the frecuency domain, and viceversa.
Unfortunately, even when we expect that this improved feedback supression algorithm is going to perform better than the previous one, the computation of the filter weights \(\mathbf h\) requires emitting impulses that can be heard by the user. Furthermore, while the acoustic response of the environment is being analyzed, the near-end should remain silent; otherwise, its own voice would be treated as an echo and attenuated by the filter.
Lets recap. Our problem is that, without a headset, we have that the signal that we send is
where \(\tilde {\mathbf {n}}\) is the version of our voice captured by our mic, and \(\hat {\mathbf {s}}\) is a aproximated-and-multy-echo version of \(\mathbf {s}\), the signal played by our speaker(s). And if we were able to make \(\hat {\mathbf {s}}=\mathbf {0}\), our feedback problem would vanish, ... at least, theoretically.
But wait ..., as a polite person, if I don’t speak when I am listening to (my interlocutor), I could assume that \(\hat {\mathbf {s}}=\mathbf {0}\) when I am speaking because my interlocutor should do the same (be silent when I speak)! But, if this is not true? A solution is: if I am speaking, can attenuate \(\mathbf {s}\) (the signal played by my speaker).
Summarizing: you should be able to control the volume of \(s\) (the analog signal that comes out from my speaker) and cut it down (or even off, if necessary) when you are speaking or your mic is recording a signal \(m\) with enough energy.