• The move to the new server is done. There are some software and database maintenance updates in process. This has us passing the hat around to help out. We appreciate any donations. Seriously, even a dollar helps. The payment page may be found here - https://www.audiokarma.org/support.html

MarmotDavid

New Member
I've been pondering this question recently - "Where does the "soundstage" that we claim to hear come from?

In practically every review of equipment, the reviewer claims that the reviewed equipment has a certain soundstage, be it wide, tall deep, flat, forward, diffuse, or whatever other adjective is used. I hear it as well when I listen to my system. Where is this coming from? I'm not asking this rhetorically; I'm asking about it technically!

If we are listening to two channel stereo sources, all the information coded into the CD, vinyl pressing, or bit stream is essentially two channels (left, right) of musical/audio information. That's it. I am speaking only here of two channel sources (not multi-channel recording, Dolby, THX etc). There is no discrete spatial information in these two channels of information other than that information can be subjectively interpreted in our brain from left/right differences. So, for example, if you had, say a violin playing on a stage to the right of center, and say you had microphones equally staggered along the horizontal plane of the stage, then the final two-channel recoding would have the right channel signal from the violin louder than the left. Our ears and brains would interpret a left/right soundstage and place the violin to the right of center. That makes sense to me.

But what makes no sense to me is other descriptions of soundstage by reviewers of equipment as "tall, deep, flat", etc., any other dimension other than the left/right horizontal dimension as attributed to the equipment itself. The only explanation, simple as it may be, is that the other dimensions of "soundstage" that we do actually hear come entirely from our own subjective interpretation of the two channel sources. These two other dimensions of soundstage, front to back and low to high, are interpreted entirely in our own heads. My hypothesis is that we certainly hear them, but I cannot imagine how we can attribute them to the components in our systems handling two channels.

In a recording session there are many microphones, and they pick up subtle spatial cues in all directions, including height and depth. But these cues are not Cartesian spatial coordinates, but rather variations in volume, pitch and tone in the final mix down to two channels, and that information ends up either in the right or left channel. Let's say that this full "indirect" the full musical spatial information is locked into those two channels, to be interpreted by a sophisticated listener in playback as spatial information. I think that is true.

So, in summary, what I am struggling with is how reviewers can attribute the front/back and low/high spatial qualities to equipment that is only handling two channel information. To me, use of the concept of "soundstage" by the equipment reviewer is sort of begging for a more measurable aspect of equipment review. Isn't the interpretation of that information happening in our subjective hearing of the audio cues? If a reviewer says that a particular amplifier, for example, has a wider soundstage, isn't that really just that the amplifier passes more of the indirect information to the listener. If that is the case, then shouldn't there be a better way to describe this aspect of two channel equipment by looking at the subtle ways that equipment handles left/right channel signal and accuracy of their handling of the subtle source information embedded in the two channels than attributing soundstage to the equipment itself? Does anyone know if three-dimensional spatial cue interpretation of two channel systems has been studied?
 
Register to hide this ad
A friend of mine dabbles with recording and mixing in a little recording studio he assembled. Nothing particularly high tech or expensive, but he has sufficient hardware and software to create some very nice sounding recordings (I'm a drummer and have recorded with him on multiple occasions). I asked him how he went about creating a multi-dimensional sound stage in his recordings (they demonstrated depth, width and height that I could reproduce on my 2-channel rig that was exactly what I "saw" sitting at his mixing station in front of Focal monitors), and he went into a technical explanation that I can't regurgitate verbatim, but the thing I remember that was primarily responsible for creating dimensionality in the recording was "by making adjustments to panning and phase during the mixing process".
 
Last edited:
I'm guessing that you don't hear a soundstage, but hear imaging.

A soundstage is not measurable. Imaging, which is left to right spread of instruments is easy to hear and most everyone hears imaging. Hearing a soundstage is mostly a product of the recording, although some gear does a better job at reproducing it.

99% of all rock pop recordings do not have a soundstage. That's because there isn't one to reproduce. The best examples of a soundstage are found in just about every direct to disk LP or audiophile-oriented recordings.

What makes them different is, all of the musicians are in the venue at the same time and playing as a cohesive unit. The recording itself is done using a minimal number of microphones (sometimes only two) and recorded live using a passive mixer to either a cutting head (for LP's) or a two-track master (for CD's). No studio manipulation is done to the "master".

IME all direct to disk LP's, many classical recordings (especially Telarc) and those from smaller audiophile-oriented labels have a soundstage. If direct to disk LP's and classical recordings are out of the question try one of the labels known for minimalist recordings. Groove Note is one of those labels for jazz.

If pop/rock is your primary type of music, it's likely you'll never hear a soundstage. It has to be in the original recording. Despite multiple attempts to artificially create one in a studio most recordings IME simply don't have one. You can't re-create what was never there.
 
Last edited:
These two other dimensions of soundstage, front to back and low to high, are interpreted entirely in our own heads. My hypothesis is that we certainly hear them, but I cannot imagine how we can attribute them to the components in our systems handling two channels.
I agree with your first sentence, but wonder why the second is difficult for you. I only occasionally get a sense of height, but depth is obvious, to me, on some recordings. Cues to depth are either present or absent in a recording, and equipment preserves those cues, or doesn't, to varying degrees.

Since we can't separate the cues from the rest of the music, they are hard to measure, and since nobody can measure what's going on in our heads (yet), we find it difficult to correlate what we hear with anything measurable. And yet, our ears and brain can hear the difference. It's a mystery.

As for myself, as Greg Brown sings, "I think I'll just - let the mystery be."
 
The Chesky Jazz Sampler Vol. 1 has an up and over test. It starts low on one side slowly gains height until it's higher than my speakers then moves to the middle then over the other speaker and slowly lowers the height. There are a couple of other sound staging tests on it also.

As for creating a soundstage by manipulating the signal in a studio, IME that generally doesn't work.
 
I agree with your first sentence, but wonder why the second is difficult for you. I only occasionally get a sense of height, but depth is obvious, to me, on some recordings. Cues to depth are either present or absent in a recording, and equipment preserves those cues, or doesn't, to varying degrees.

Since we can't separate the cues from the rest of the music, they are hard to measure, and since nobody can measure what's going on in our heads (yet), we find it difficult to correlate what we hear with anything measurable. And yet, our ears and brain can hear the difference. It's a mystery.

As for myself, as Greg Brown sings, "I think I'll just - let the mystery be."

Credit to Iris DeMent for "let the mystery be".
 
Back when, we paid attention to an amp’s channel separation spec. Wouldn’t that be the objective parameter defining the Equipment’s ablity to affect “soundstage”.
Then, wouldn’t all other nuances of soundstage be a function of the production process?
I think talking about a piece of equipment having this or that amount of soundstage is jibber-jabber
But, hey, that’s just me; no offense intended.
 
I've been pondering this question recently - "Where does the "soundstage" that we claim to hear come from?

In practically every review of equipment, the reviewer claims that the reviewed equipment has a certain soundstage, be it wide, tall deep, flat, forward, diffuse, or whatever other adjective is used. I hear it as well when I listen to my system. Where is this coming from? I'm not asking this rhetorically; I'm asking about it technically!

If we are listening to two channel stereo sources, all the information coded into the CD, vinyl pressing, or bit stream is essentially two channels (left, right) of musical/audio information. That's it. I am speaking only here of two channel sources (not multi-channel recording, Dolby, THX etc). There is no discrete spatial information in these two channels of information other than that information can be subjectively interpreted in our brain from left/right differences. So, for example, if you had, say a violin playing on a stage to the right of center, and say you had microphones equally staggered along the horizontal plane of the stage, then the final two-channel recoding would have the right channel signal from the violin louder than the left. Our ears and brains would interpret a left/right soundstage and place the violin to the right of center. That makes sense to me.

But what makes no sense to me is other descriptions of soundstage by reviewers of equipment as "tall, deep, flat", etc., any other dimension other than the left/right horizontal dimension as attributed to the equipment itself. The only explanation, simple as it may be, is that the other dimensions of "soundstage" that we do actually hear come entirely from our own subjective interpretation of the two channel sources. These two other dimensions of soundstage, front to back and low to high, are interpreted entirely in our own heads. My hypothesis is that we certainly hear them, but I cannot imagine how we can attribute them to the components in our systems handling two channels.

In a recording session there are many microphones, and they pick up subtle spatial cues in all directions, including height and depth. But these cues are not Cartesian spatial coordinates, but rather variations in volume, pitch and tone in the final mix down to two channels, and that information ends up either in the right or left channel. Let's say that this full "indirect" the full musical spatial information is locked into those two channels, to be interpreted by a sophisticated listener in playback as spatial information. I think that is true.

So, in summary, what I am struggling with is how reviewers can attribute the front/back and low/high spatial qualities to equipment that is only handling two channel information. To me, use of the concept of "soundstage" by the equipment reviewer is sort of begging for a more measurable aspect of equipment review. Isn't the interpretation of that information happening in our subjective hearing of the audio cues? If a reviewer says that a particular amplifier, for example, has a wider soundstage, isn't that really just that the amplifier passes more of the indirect information to the listener. If that is the case, then shouldn't there be a better way to describe this aspect of two channel equipment by looking at the subtle ways that equipment handles left/right channel signal and accuracy of their handling of the subtle source information embedded in the two channels than attributing soundstage to the equipment itself? Does anyone know if three-dimensional spatial cue interpretation of two channel systems has been studied?

You aren't looking at the entire situation. Imagine a stereo microphone array in a room. You have a vocalist standing at the microphone, a bass player to the left and behind her, there's a guitar player on the right, and a drummer behind her. This will be a minimalist recording.

The singer, being right at the microphone will have minimal interaction with the room. The bass player will have much more room interaction, as will the guitar, both bouncing energy off of their respective side walls. The drummer will interact with the rear walls and the side walls. Your brain will naturally triangulate a sound origination point based on the position from the microphone.

To make this happen you need good equipment, properly set up, and a recording with the needed information. My go to is the Chesky recording of "Johnny Frigo with John and Bucky Pizzarelli. Their first recording and I think by far the best at soundstage.
 
So, in summary, what I am struggling with is how reviewers can attribute the front/back and low/high spatial qualities to equipment that is only handling two channel information.
Similarly, you have but two ears. Let's start out in the world around us.

When I'm sitting in my back yard, the "soundstage" is decidedly three dimensional. I hear dogs barking from neighbor's yards in most every direction. There are cues I detect where my ears are able to tell me both the direction and distance. I hear an airplane flying in the area. I not only detect direction and distance, but height as well. And can track its progress across the sky even if I cannot see it.

Lets get closer with a musical example. Ever attend a symphony performance? Close your eyes and you can similarly determine not only direction but depth as well. Percussion and brass at the back of the orchestra. Triangles seemingly hang in the air. Woodwinds somewhat closer with strings in front.

Good recordings capture those cues we hear in the natural world and gear with exceptional transparency is able to convey it. As Joe indicated, such is all but absent is most close miked, multi-tracked recordings. Simple panning can place instruments and voice left to right, but are unable to truly duplicate the sense of depth with mere knob twiddling with a recording console. I really enjoy music by Dead Can Dance but most of their recordings are flat as a board. Painted ships on a painted sea.

Ever look through a GAF viewer? You know those things with two images for each scene? "Wow! Look at the Grand Canyon! I can see how deep it is and how far it stretches." Most recordings use a very different approach. Imagine doing what the mikes are. Take a picture of a singer's mouth. I'm not joking. Just visualize where a typical microphone is with respect to the performer. Add other close up pictures of various instruments. Make slides out of these pictures are hold them up together and look through the light. Are you able to determine relationships between these instruments? Of course not!

So, one must begin with a recording capable of conveying depth. Which usually requires a minimally miked approach that captures these sound relationships like our ears do. Minimal miking is more distant and picks up not only the instruments but the natural hall ambience as well. They are then played back not terribly dissimilar to that GAF viewer example. Wow! You can hear those relationships that convey both location and depth. Then you are able to evaluate what audio gear is capable of reproducing that which is found in the recording. Some of the best examples in my library are found on Telarc, Reference Recordings and Windham Hill labels among others.
 
Last edited:
E-Stat did a great job of drawing the picture of sound around you and the similarity to listening to pre-recorded sound playback. I'd simply add that the simple stereo mic and then the simple stereo speaker recording playback is an added process in the middle. And while we certainly lose with it, when well done it can retain much of the original 'scene'. The better then entire 'chain' (both halves...recording studio and the playback venue), the more realistic it can be.

When it's well done, I like to think of looking out through the mic opening and 'seeing' the performance, with my ears able to hear what the two mics are hearing.

Pan pot sweeping in the recording console may place objects somewhere between the two speakers. Indeed, as you pan you can follow the sound in an arc around from side to side. But it's really not at all the same as the much superior recording naturally with two mics process. The reason why is that distance from the performer to the mic is directly tied to phase and to loudness. And the difference of those between the two mics being very key. When you move a pan pot from side to side you change the level difference between channels, but it cannot change the timing to match the place the sound is now coming from.

When you want a real sounding stage presence with distance appropriate between the performers, then the vastly better representation will come from it being naturally mic'ed and not artificially contrived. Folded, spindled and mutilated is the phase that always comes to mind for me.

EV3
 
And the difference of those between the two mics being very key. When you move a pan pot from side to side you change the level difference between channels, but it cannot change the timing to match the place the sound is now coming from.
Let me just clarify that by "minimal miking", I don't always refer to using only two. The wonderful unprocessed Windham Hill recordings of George Winston, Michael Hedges, Liz Story. et. al. do indicate two mics were used. The Telarc recording of The Firebird, however, used five: three in front with two further back in the orchestra behind the perch I took for timing. It was Jack Renner's evolution of the three mic approach championed by Wilma Fine-Cozart for the wonderful Mercury Living Presence recordings in the 50s. I forgot to mention those as well as I have a couple of the stereo mix downs. They were also released in multi-channel SACD preserving the original front three.
 
Hi all- having written the opening post, I'd like to redirect the discussion because I think it got off track a bit. Thanks for all the comments up to here. If there was any misunderstanding about my opening post, I do hear the imaging and soundstage and agree with you all that it is embedded information in the stereo tracks creation process for all the reasons that were nicely explained by several members. I said that information in the stereo tracks wasn't discrete (like a Dolby track).

Thud3 captured my struggle exactly. Thank you, Thud 3, for saying it more directly than I did. He said: "Back when, we paid attention to an amp’s channel separation spec. Wouldn’t that be the objective parameter defining the Equipment’s ablity to affect “soundstage”. Then, wouldn’t all other nuances of soundstage be a function of the production process? I think talking about a piece of equipment having this or that amount of soundstage is jibber-jabber. But, hey, that’s just me; no offense intended."

My issue/struggle is that reviewers of equipment continually talk about a given piece of equipment's "soundstage". If the information for a resulting soundstage in our subjective perception is only contained in the original left/right tracks of audio, then why is it that reviewers are talking about the soundstages of cartridges, DACs, preamps and amplifiers and speakers? And how in the equipment chain do they know which piece of equipment is causing the so-called soundstage of the equipment? If all that is really happening (which I suspect) is that some equipment passes the complex audio signal and sounds better to the end point (your brain), then we should be talking about signal integrity of stereo separation or something through each equipment piece.

In summary, the soundstage or sound imaging if you like that term better, is the subjective result in your mind of the playing of well recorded stereo tracks, is it not? Yes, of course the sound engineers do their best to embed the spatial cues in those tracks with sophisticated miicrophone placement and mixing. But why are we talking about equipment having soundstages?
 
E-stat got it...some equipment is simply better than others at retrieving that information in the recording that deals with spatial cues.
Is it bothering you that there's no objective metric for this, @MarmotDavid?
That's what gets me about the measurements crowd (particularly over at ASR)...they seem to be totally oblivious to the relevance of image and sound stage resolution in the reproduction of a recording. It's as if it won't even be acknowledged if it can't be measured...and that's where I can't get on board with the measurements crowd.
 
E-stat got it...some equipment is simply better than others at retrieving that information in the recording that deals with spatial cues.
Is it bothering you that there's no objective metric for this, @MarmotDavid?
That's what gets me about the measurements crowd (particularly over at ASR)...they seem to be totally oblivious to the relevance of image and sound stage resolution in the reproduction of a recording. It's as if it won't even be acknowledged if it can't be measured...and that's where I can't get on board with the measurements crowd.
Perhaps small unmeasurable effects change or delay or distort the spatial cues? That would affect the perception of a stage.
 
Perhaps small unmeasurable effects change or delay or distort the spatial cues? That would affect the perception of a stage.
That certainly is the case when acoustic issues from the room interfere with a faithful reproduction of what's in the recording...spatial cues in particular.
 
Back
Top Bottom