• The move to the new server is done. There are some software and database maintenance updates in process. This has us passing the hat around to help out. We appreciate any donations. Seriously, even a dollar helps. The payment page may be found here - https://www.audiokarma.org/support.html

Need expert advice on mitigating the harmful effects of digital/compressed sound (mostly voices)

TerpSound

New Member
I'm a professional conference and business interpreter, and some of my work is now remote/online. The sound I have to work with is often compressed and not professional, and at times simply poor. I rarely have the chance to advise participants what equipment and settings to use (e.g. on Zoom).
Since poor ("toxic") sound can cause hearing damage even at lower volumes, I need to find a solution that is both practical and absolutely safe for long hours of listening.

Here's a short(ish) article that describes the damage that digital/compressed sound can do:
https://aiic.org/uploaded/web/UN%20NegDel/Technical%20Note%20to%20the%20CEB.pdf
More resources here (see the section on remote simultaneous interpreting /RSI/):

The sound needs to get from a Windows laptop to my Sennheiser HD 25, and the sound from my mic (currently a Sennheiser USB-C lavalier, but could get another) needs to get back to the laptop. There needs to be tone control for improved intelligibility, preferably bass, mids and treble, and it needs to be physical knobs (or perhaps faders). Since intepreting requires intense focus and is very exhausting even under ideal conditions, and I'm using two screens as it is, only automatic software EQ is realistic, and tone control must be simple and convenient enough to manage without much attention.

Solutions so far considered:
Software EQ - for calibration to the headphones and potentially mitigating some of the damaging aspects of sound
A decent DAC (e.g. SMSL)
Vintage integrated amp (any brand with the right tone control and quality up to about 400 EUR)
Headphone amp - none have tone control these days
DJ mixer with or without USB interface - most small mixers do not have the functions I need, the bigger ones are an overkill, and there's no need for mixing
Audio interface (e.g. Focusrite) - it does not have tone control
MIDI controller - great for tone control, but no use outside of work, potentially expensive and/or too big
Small and cheap DAC/AMP combos with tone control - potentially not good enough sound, unreliable
AaptEar, an in-line limiter by LimitEar (UK) - company sells at their discretion, no webshop,

Any ideas, recommendations are welcome. Some compromise will be necessary it seems, but it should be reasonable.
 
Register to hide this ad
You might be better off consulting an audiologist for advice as opposed to us.

Since the main recommendation of the PDF you linked is to have fewer assignments for interpreters, therefore limiting exposure, I don’t think you can fix the issue with more devices because the “toxic” artifacts are baked in to what you receive, so to speak.

The technical advice boils down to: use the “HiFi” mode in Zoom and WebEx.

I’m not aware of any devices that restore dynamic range after compression has been applied, for example. I may be wrong. Your gear might already be good enough, since it is not of the “cheap” variety the PDF mentions (though their lack of detail on this score is frustrating).

This is fascinating though. A conversation with an audiologist would be worthwhile.
 
I'd be considering turning all of the audio into text. Don't listen to any of it.
 
Thanks for responding!
At the moment, my equipment consists of the above mentioned headphones and mic, plus Equalizer APO+Peace and some common sense. All the other devices are ones that I researched and considered, and could potentially use.
And I haven't had a lot of remote work partly because I knew it's risky, and as a freelancer I have to take care of my own health and safety. There are ISO standards in place, but I have no way of enforcing them, clearly. They apply to employers of interpreters like the UN, the EU and NATO, and theoretically to interpreting platforms established for the sole purpose of interpreting services.
But I'm a pro service provider, so I want to offer my clients the option of remote interpreting. I have a client now who sometimes book me for in-person assignments, and sometimes for remote.

Thing is, my use case lies at the meeting point of various areas of expertise, and none of the solutions in any are a perfect fit, and experts in any of these areas find it hard to see our needs (as interpreters) clearly.
It is recommended for us to regularly see an audiologist, but they would never be able to advise on technical aspects.It's hard for me to imagine that there is simply no technical solution(s) at all. There must be something I'm not seeing...
Until I find something though, I will not accept certain jobs.
 
I'd be considering turning all of the audio into text. Don't listen to any of it.
It is technically impossible. AIs can only do so much. There are speech to text applications, and some people use them, but they are very imperfect. Completely unreliable for professional purposes.
And there is the not insignificant issue, that sight translation (to give it its proper name) is not any less demanding, and it strains my eyes. I couldn't possibly do it for hours on end. It also potentially slows me down, which is not an option.
 
I have no idea how you can fixed compressed sound. Also don't see how low volume compressed audio damages hearing......new one on me. But what I don't know could fill a set of encyclopedias.

Worked in industrial settings all my career and the rule was hearing protection if the ambient noise level was higher than 85dB. Higher SPLs damage hearing for sure, the higher the dB the quicker the damage.
 
I have no idea how you can fixed compressed sound. Also don't see how low volume compressed audio damages hearing......new one on me. But what I don't know could fill a set of encyclopedias.

Worked in industrial settings all my career and the rule was hearing protection if the ambient noise level was higher than 85dB. Higher SPLs damage hearing for sure, the higher the dB the quicker the damage.
The article above gives you an idea about the issues with digital sound in general.
I found this explanation with a quick google search in a LinkedIn post by an interpreter (many of us have been posting and looking for solutions for a long time now). Not very technical, but gives you an idea:
"one of the biggest problems we have with sound in RSI is that its range is compressed, and at the same time you lose certain frequencies altogether, while the remaining ones are boosted unnaturally. This means that much of the energy of the sound (up to 70%) is delivered in a narrow frequency band at around 4kHz, which can be damaging even if your volume is within safe limits, even within a relatively short time."

So it's not just volume that matters.

Software EQ can address the excessive volume of specific frequencies, so that's great, but it does not redistribute energy naturally across all frequencies.
 
maybe the audio could be truncated somehow below 4000 Hz
 

Attachments

  • IMG_9837.png
    IMG_9837.png
    193.7 KB · Views: 6
maybe the audio could be truncated somehow below 4000 Hz
My approach - without being an expert on this - would be some form of complex EQ, perhaps using AI (at very, very low latency) that would take the amount of excess volume on those overloaded frequencies and would redistribute it across all frequencies using some general metrics for male and female voices as a template. This could theoretically produce a very natural result.
Another technically pretty easy but hardly explored approach could be to collect some EQ/frequency data about the speaker"s voice at their end, and use it the other end to recreate their sound. Not sure it's doable or whether it's legally acceptable.
And of course, eg Zoom itself could also simply send every listener some info on the algorithm used to compress speakers' voices, so that it is then used to undo the compression.
 
The article above gives you an idea about the issues with digital sound in general.
I guess this is tangential to your query, but I don't interpret (no pun intended) this as an indictment of digital sound as a whole. In fact one of the examples in the article is dentist's drills, an analog device.

The danger is constant ingestion of a narrow frequency, something that's not exclusive to the digital domain. I see it as something akin to carpal tunnel syndrome; damage to our organs by their use in a repetitive motion that evolution did not forsee.

As for the solution unfortunately I cannot contribute, sorry! There are many problems that generative AI can likely solve in the coming years, audio processing being among them. However the greater issue on this front may be one being faced by many skilled professionals in that the very technology assisting them at the current time may make their very vocations obsolete, and I'd not count interpreters as being immune to this issue.
 
I guess this is tangential to your query, but I don't interpret (no pun intended) this as an indictment of digital sound as a whole. In fact one of the examples in the article is dentist's drills, an analog device.

The danger is constant ingestion of a narrow frequency, something that's not exclusive to the digital domain. I see it as something akin to carpal tunnel syndrome; damage to our organs by their use in a repetitive motion that evolution did not forsee.

As for the solution unfortunately I cannot contribute, sorry! There are many problems that generative AI can likely solve in the coming years, audio processing being among them. However the greater issue on this front may be one being faced by many skilled professionals in that the very technology assisting them at the current time may make their very vocations obsolete, and I'd not count interpreters as being immune to this issue.
There is already Ai interpreting, various providers have made various very optimistic promises, but they have failed to resolve one very important underlying issue: AIs do not UNDERSTAND language. They just generate the mathematically most likely translation. They "hallucinate", which is a glorified way of saying they simply don't work properly. What AI interpreting there is, is flawed ti the point of being completely useless. there are stories circulating about clients entrusting an important event to AI interpreting provider only then to end up calling every human interpreter they can run down in a mad rush and beg them to save their "important" event.

When AIs start understanding language exactly like humans, interpreters MAY be replaces in some applications. One thing an AI will never have is name accountability. Where that's a requirement, AIs will not replace humans.

I really do think however that AI audio processing is just round the corner, if not already here...
 
There is already Ai interpreting, various providers have made various very optimistic promises, but they have failed to resolve one very important underlying issue: AIs do not UNDERSTAND language. They just generate the mathematically most likely translation. They "hallucinate", which is a glorified way of saying they simply don't work properly. What AI interpreting there is, is flawed ti the point of being completely useless. there are stories circulating about clients entrusting an important event to AI interpreting provider only then to end up calling every human interpreter they can run down in a mad rush and beg them to save their "important" event.

When AIs start understanding language exactly like humans, interpreters MAY be replaces in some applications. One thing an AI will never have is name accountability. Where that's a requirement, AIs will not replace humans.

I really do think however that AI audio processing is just round the corner, if not already here...
Understood, thanks for the info. It is invading my [completely different] industry as well for better or for worse. I hope you are right and you remain gainfully employed (sans hearing damage of course).
 
I checked it a while ago, and I don't see any product there that could be useful and tick all the (actually not very numerous) boxes.

Qudelix is more advanced, but even they don' have what interpreters need.
 
Generative AI will eventually replace interpreters. Many, many, many other human professions also.
 
Understood, thanks for the info. It is invading my [completely different] industry as well for better or for worse. I hope you are right and you remain gainfully employed (sans hearing damage of course).
At the moment, I have more work than in quite a while, both translation and interpreting. There was a lull when some clients were swayed by AI promises, but they are now back.
Also, what I do is quite rare and I'm also more versatile than most, so I have hope.
There is also a worrying trend of younger generations raised on screens and social media (all their lives or for part of it) missing key basic skills to join the translation or interpreting profession. So I won't have a lot of competition it seems (strictly speaking I don't have a lot of competition even now, but I also have to convince clients of that).
 
Generative AI will eventually replace interpreters. Many, many, many other human professions also.
Not in professional applications, no. Not in court, where you have legal responsibility and need accountability. What will you do if an AI perjures itself and you suffer the consequences?
Same in many other applications. AIs don't understand language. they are faking it.
When AIs really start understanding language and become sentient, then yes, interpreters MAY be replaced in some more applications. But if that happens, intepreting will be the least of our problems.

At the moment, AIs are very limited. They can barely manage the simplest literal translations in writing, let alone a complex job in simultaneous interpreting. And they are built on plagiarism and IP theft.
 
AK member @Cask05 has developed demastering techniques to restore dynamic range to recorded music, but ultimately, it's a challenge to de-compress recordings.
...and I'd have to do that on the fly with low latency. Without much drain on the CPU, as I already have a lot of applications running.
Using EQ is not enough, I'm doing that already. It can only reduce/boost the volume of specific frequencies (apply a curve), but it does not recreate anything. I'm not even sure Equalizer APO is intelligent enough not to just reduce volume by a certain amount but keep it at a specific level (with a limiter in place on each selected frequency).
And all the while, the voices/speech need to remain easily understandable and clear, not just safe.

At this point I wouldn't even care if the speakers' voices were significantly altered as long as they are safe and clear. They'd have to sound natural to avoid excessive fatigue, but they can be altered.
 
Here is the latest in my bag of tricks using multiband expander/compressor and dynamic EQ plugins to regain audio track dynamic range of favorite tracks--those damaged during mastering. There are two such plugins available from FabFilter Pro (with 1 month free trial for evaluation):

Pro-MB (for "multiband expander"), and
Pro-Q4 (for "dynamic EQ")

I find them extremely easy and powerful to use--since these are the same plugins used by the mastering guys to compress the tracks in the first place. Reversing those effects using the same (professional) plugins is the closest thing to a real solution that anyone is ever going to find (IMHO).

However, you must use your eyes and ears to use these plugins to reverse mastering compression, limiting, and EQ. It's pretty straightforward once you can see the dynamics of each audio track moving around during real-time playback, and see the limiting (clipping) within Audacity. Static and dynamic EQ can be applied to find the most effective reverse EQ settings for each track, then save the results to HD for future playback (instead of the tracks found on the compact disc or original download tracks).

Without much drain on the CPU, as I already have a lot of applications running.

You can run these plugins on something like Audacity and save the tracks in order to avoid your self-imposed issue of "no added computational load". This is what I do for the music tracks I listen to often but are otherwise compressed (damaged) to the point of needing something to reverse the damage. Then you only need to do this once, and then you can forget about having to recall the particular settings for each track in the future.

However, note that it is equally easy to run these plugins in real time during playback, as I find their added computational load to be quite low compared to what I'd otherwise guess they would be.

Chris
 
Last edited:
Back
Top Bottom