I’d really like to see an improved key change algorithm in VirtualDJ that includes proper formant correction.
At the moment, changing the key of a track can make vocals sound noticeably unnatural, particularly when moving more than one or two semitones. The pitch changes, but the vocal character shifts with it, which can produce the familiar “chipmunk” or overly deep effect (formant frequencies).
DAW and other audio software systems can adjust pitch while preserving the formants; this makes the result sound far more natural. It would be extremely useful to have something similar available in VirtualDJ, either as an enhanced option within Key Match/Key Cue or as a separate high-quality pitch mode.
Ideally, it could work on either:
The entire track
The vocal stem only
Individual stems where appropriate
Applying formant correction specifically to the vocal stem may also reduce processing demands while producing a better overall result. This would make harmonic mixing much more flexible, especially when a track needs to be moved by several semitones to fit properly. It would also make the Key Cue pads far more usable for creative live performance, rather than key changes being limited by how unnatural the vocals begin to sound.
Is this something that could be considered for a future update?
At the moment, changing the key of a track can make vocals sound noticeably unnatural, particularly when moving more than one or two semitones. The pitch changes, but the vocal character shifts with it, which can produce the familiar “chipmunk” or overly deep effect (formant frequencies).
DAW and other audio software systems can adjust pitch while preserving the formants; this makes the result sound far more natural. It would be extremely useful to have something similar available in VirtualDJ, either as an enhanced option within Key Match/Key Cue or as a separate high-quality pitch mode.
Ideally, it could work on either:
The entire track
The vocal stem only
Individual stems where appropriate
Applying formant correction specifically to the vocal stem may also reduce processing demands while producing a better overall result. This would make harmonic mixing much more flexible, especially when a track needs to be moved by several semitones to fit properly. It would also make the Key Cue pads far more usable for creative live performance, rather than key changes being limited by how unnatural the vocals begin to sound.
Is this something that could be considered for a future update?
geposted Sat 25 Jul 26 @ 10:08 am
I'm not sure I understand the request..
If you want to change pitch but keep key the same, that's the job of "master tempo" and "keylock" (which are two similar but different options / settings that already exist)
If you want to change the key only on a stems part that would be interesting, but it would clash with the rest stems (unless you isolate the shifted stem)
What I mean is that if the key of a track is C and you only shift vocals one key up/down, the vocals will clash harmonically with the instrumental part of the same song.
So, what are you actually wanting to do ?
If you want to change pitch but keep key the same, that's the job of "master tempo" and "keylock" (which are two similar but different options / settings that already exist)
If you want to change the key only on a stems part that would be interesting, but it would clash with the rest stems (unless you isolate the shifted stem)
What I mean is that if the key of a track is C and you only shift vocals one key up/down, the vocals will clash harmonically with the instrumental part of the same song.
So, what are you actually wanting to do ?
geposted Sun 26 Jul 26 @ 6:48 am
On certain types of vocal effects processors (like pitch correction or vocoders) there's often a control which is able to adjust the sound of the voice being processed, to make it sound more male or more female, without changing the pitch.
From Google:
"Formant shifting is an audio processing technique that alters the tone or timbre of a sound—particularly a human voice—without changing its fundamental pitch. By shifting the harmonic resonances (formants), it simulates changes in the size of the vocal tract, making voices sound deeper, higher, more masculine, or more feminine"
From Google:
"Formant shifting is an audio processing technique that alters the tone or timbre of a sound—particularly a human voice—without changing its fundamental pitch. By shifting the harmonic resonances (formants), it simulates changes in the size of the vocal tract, making voices sound deeper, higher, more masculine, or more feminine"
geposted Sun 26 Jul 26 @ 9:05 am
Thanks for the replies. I think there may be a slight misunderstanding about what I mean by formant correction.
I am not asking for the vocal stem to be transposed into a different musical key from the instrumental. The whole track would still be shifted by the same number of semitones, so the vocals and instruments would remain harmonically aligned.
The issue is that a vocal contains two different things:
Fundamental pitch, which determines the musical note being sung
Formants, which are resonant frequencies created by the shape and size of the singer’s vocal tract
When an ordinary pitch shifting algorithm moves a vocal upwards, it often moves both the musical pitch and the formants upwards together. The notes are technically correct, but the singer’s vocal character becomes thinner, smaller and more “chipmunk-like”.
Moving the pitch down can have the opposite effect, making the voice sound unnaturally deep, heavy or “giant-like”.
Formant correction separates those two elements. The musical pitch is changed to the new key, but the vocal resonances are compensated so that the singer retains more of their original vocal character.
For example:
Original vocal
Pitch: 0 semitones
Formants: unchanged
Ordinary pitch shift
Pitch: +3 semitones
Formants: also shifted upwards
Result: correct musical key, but the voice sounds noticeably smaller and less natural
Pitch shift with formant correction
Pitch: +3 semitones
Formants: kept close to their original position
Result: correct musical key, with a much more natural sounding voice
This is different from Master Tempo or Keylock. Those features allow the tempo to change while trying to preserve the original musical key. I am suggesting improved processing for the opposite situation: deliberately changing the musical key while preserving the natural vocal timbre as far as possible.
Applying it only to the vocal stem was simply one possible way of reducing processing demands. The instrumental stems would still be transposed by the same amount, but they generally do not need vocal formant correction.
Cubase and various other pitch processing systems already provide separate pitch and formant controls. I am suggesting similar compensation could improve VirtualDJ’s Key Match and Key Cue results, particularly when moving a song by several semitones.
Here is a diagram that attempts to explain formants visually.
I am not asking for the vocal stem to be transposed into a different musical key from the instrumental. The whole track would still be shifted by the same number of semitones, so the vocals and instruments would remain harmonically aligned.
The issue is that a vocal contains two different things:
Fundamental pitch, which determines the musical note being sung
Formants, which are resonant frequencies created by the shape and size of the singer’s vocal tract
When an ordinary pitch shifting algorithm moves a vocal upwards, it often moves both the musical pitch and the formants upwards together. The notes are technically correct, but the singer’s vocal character becomes thinner, smaller and more “chipmunk-like”.
Moving the pitch down can have the opposite effect, making the voice sound unnaturally deep, heavy or “giant-like”.
Formant correction separates those two elements. The musical pitch is changed to the new key, but the vocal resonances are compensated so that the singer retains more of their original vocal character.
For example:
Original vocal
Pitch: 0 semitones
Formants: unchanged
Ordinary pitch shift
Pitch: +3 semitones
Formants: also shifted upwards
Result: correct musical key, but the voice sounds noticeably smaller and less natural
Pitch shift with formant correction
Pitch: +3 semitones
Formants: kept close to their original position
Result: correct musical key, with a much more natural sounding voice
This is different from Master Tempo or Keylock. Those features allow the tempo to change while trying to preserve the original musical key. I am suggesting improved processing for the opposite situation: deliberately changing the musical key while preserving the natural vocal timbre as far as possible.
Applying it only to the vocal stem was simply one possible way of reducing processing demands. The instrumental stems would still be transposed by the same amount, but they generally do not need vocal formant correction.
Cubase and various other pitch processing systems already provide separate pitch and formant controls. I am suggesting similar compensation could improve VirtualDJ’s Key Match and Key Cue results, particularly when moving a song by several semitones.
Here is a diagram that attempts to explain formants visually.
geposted Sun 26 Jul 26 @ 1:14 pm
I assume this infographic is AI generated and the spectrograms and waveforms shown there are garbage illustrations that don't actually have any meaning?
Other than that, I haven't looked into the details yet, but the difficult thing may be recognizing the formants in the first place. While in a DAW you probably have a separate track for each singer, VirtualDJ only has access to all vocals together, which could include multiple singers, even at the same time.
Other than that, I haven't looked into the details yet, but the difficult thing may be recognizing the formants in the first place. While in a DAW you probably have a separate track for each singer, VirtualDJ only has access to all vocals together, which could include multiple singers, even at the same time.
geposted Sun 26 Jul 26 @ 1:22 pm
Yes, the infographic was made as a simple visual explanation. And just to show what waveforms and spectrograms look like, those weren't from a real recording, so please ignore those particular graphics.
What I was trying to say is, still the written explanation. In a program such as Spectralayers (Steinberg), you can identify the formants and from a vocal and preserve or not preserve them when applying stretch and pitch functions.
I agree that a full commercial recording, or even an extracted vocal stem with several singers, is more difficult to process than an isolated monophonic vocal. Overlapping singers, harmonies, reverb and stem separation artefacts would make accurate processing more difficult.
However, I don't think the system would have to identify each singer individually or explicitly recognise a fixed set of formant frequencies for each voice necessarily.
As far as I know, pitch shifting that preserves formants can be done by estimating the spectral envelope or resonant character of the signal over time, shifting the underlying pitch and then compensating the spectral envelope such that it doesn't just go up or down by the same amount.
In other words this:
Pitch shift only
Original pitch structure -> Up 3 semitones
Original spectral envelope +3 semitones also moved
Result → musically correct, but the vocal character becomes unnaturally small or bright
While this:
Pitch shift preserving formants
Original pitch structure -> transposed up by 3 semitones
Original spectral envelope → preserved in about original frequency region
more of the original vocal character preserved, but musically correct
This is not some theoretical feature. For example, Steinberg’s pitch processing is formant-preserving, so that the formants remain where they are as the pitch changes, and you don’t get the familiar “Mickey Mouse” or “monster” effects.
WaveLab also allows you to correct the formants of vocal material and includes specific methods for monophonic, speech and multi-purpose material.
It doesn’t mean it would work flawlessly on each and every mixed or stem-separated recording. Presumably multiple singers at once would reduce accuracy . There may need to be different quality modes depending on available CPU . Thatt being said, even an imperfect formant preserving mode might sound much more natural than simply shifting the entire vocal spectrum up or down with the key.
Therefore my initial idea of placing it on the vocal stem was not about changing the vocal to another musical key. Each stem would still undergo exactly the same pitch change. The idea was simply to apply the additional formant compensation where it would be most useful, rather than expending the same processing effort on drums and other material that does not contain a human vocal tract.
So I accept that it might technically challenging to implement. The feature request is basically asking if a better quality, formant-preserving pitch mode could be investigated, not claiming it would be trivial to implement.
FYI, pay it forward, a kind word and a genuine handshake favours constructive dialogue.
What I was trying to say is, still the written explanation. In a program such as Spectralayers (Steinberg), you can identify the formants and from a vocal and preserve or not preserve them when applying stretch and pitch functions.
I agree that a full commercial recording, or even an extracted vocal stem with several singers, is more difficult to process than an isolated monophonic vocal. Overlapping singers, harmonies, reverb and stem separation artefacts would make accurate processing more difficult.
However, I don't think the system would have to identify each singer individually or explicitly recognise a fixed set of formant frequencies for each voice necessarily.
As far as I know, pitch shifting that preserves formants can be done by estimating the spectral envelope or resonant character of the signal over time, shifting the underlying pitch and then compensating the spectral envelope such that it doesn't just go up or down by the same amount.
In other words this:
Pitch shift only
Original pitch structure -> Up 3 semitones
Original spectral envelope +3 semitones also moved
Result → musically correct, but the vocal character becomes unnaturally small or bright
While this:
Pitch shift preserving formants
Original pitch structure -> transposed up by 3 semitones
Original spectral envelope → preserved in about original frequency region
more of the original vocal character preserved, but musically correct
This is not some theoretical feature. For example, Steinberg’s pitch processing is formant-preserving, so that the formants remain where they are as the pitch changes, and you don’t get the familiar “Mickey Mouse” or “monster” effects.
WaveLab also allows you to correct the formants of vocal material and includes specific methods for monophonic, speech and multi-purpose material.
It doesn’t mean it would work flawlessly on each and every mixed or stem-separated recording. Presumably multiple singers at once would reduce accuracy . There may need to be different quality modes depending on available CPU . Thatt being said, even an imperfect formant preserving mode might sound much more natural than simply shifting the entire vocal spectrum up or down with the key.
Therefore my initial idea of placing it on the vocal stem was not about changing the vocal to another musical key. Each stem would still undergo exactly the same pitch change. The idea was simply to apply the additional formant compensation where it would be most useful, rather than expending the same processing effort on drums and other material that does not contain a human vocal tract.
So I accept that it might technically challenging to implement. The feature request is basically asking if a better quality, formant-preserving pitch mode could be investigated, not claiming it would be trivial to implement.
FYI, pay it forward, a kind word and a genuine handshake favours constructive dialogue.
geposted Sun 26 Jul 26 @ 2:20 pm
I think this is key
I'm almost certain plugins like Melodyne would be acting on an isolated vocal track.
Also this sounds like something that could be a bit more intensive than the basic key shift, and may get worse with more voices involved with different tones (e.g. a choir).
A good request nonetheless. I would vote for it to be optional though - sometimes the chipmunk effect is actually wanted.
Adion wrote :
While in a DAW you probably have a separate track for each singer, VirtualDJ only has access to all vocals together, which could include multiple singers, even at the same time.
I'm almost certain plugins like Melodyne would be acting on an isolated vocal track.
Also this sounds like something that could be a bit more intensive than the basic key shift, and may get worse with more voices involved with different tones (e.g. a choir).
A good request nonetheless. I would vote for it to be optional though - sometimes the chipmunk effect is actually wanted.
geposted Sun 26 Jul 26 @ 2:40 pm
I think the main issue here would be the sheer amount of processing the vocal would be going through.
Imagine if you slow the track tempo down, but you've got key lock (master tempo) on, so the key stays the same. That's one process. Then maybe change the key up, to match another track. Process #2.
Separate the vocal stem in order to apply formant correction. Process #3.
Apply formant correction. Process #4. I'd imagine by the end of that lot, it'd sound pretty mangled.
Imagine if you slow the track tempo down, but you've got key lock (master tempo) on, so the key stays the same. That's one process. Then maybe change the key up, to match another track. Process #2.
Separate the vocal stem in order to apply formant correction. Process #3.
Apply formant correction. Process #4. I'd imagine by the end of that lot, it'd sound pretty mangled.
geposted Sun 26 Jul 26 @ 4:04 pm
Yes, as an effect, pitch shifting vocals is a well established technique for introducing character. Our ancestors were experimenting with tape speeds and synchronised tape loops long before samplers became financially viable for most musicians. Delia Derbyshire and the BBC Radiophonic Workshop are obvious examples. I suspect many of us here could reel off a sizeable list of popular records that use this technique to great effect.
However, that creative use is slightly different from what I am proposing. Sometimes the chipmunk or monster effect is desirable. At other times, particularly when changing a track’s key for harmonic mixing, the aim is to retain as much of the original vocal character as possible. That is why I agree it should be optional.
The Roland VT-4 is a dedicated hardware unit with onboard DSP. Its architecture is designed to perform a limited set of pitch and formant processes in real time with very low latency. A contemporary laptop or MacBook has considerably greater overall processing capability, although of course it is also being asked to perform many more tasks at once.
VirtualDJ already separates stems and routes selected stems through independent effect-processing paths. It is therefore reasonable to ask whether a formant-preserving stage could eventually be integrated into the vocal stem’s existing processing path.
I currently work with SpectraLayers Pro 12. I have not upgraded to version 13 yet, as Steinberg often offers upgrade deals several months after release. Other software may be preferred by different users, of course.
Different strokes and all that jazz.
SpectraLayers is a professional spectral audio editor and source-separation application. Rather than presenting audio solely as a conventional waveform, it displays sound across time, frequency and amplitude. This allows individual elements within a recording to be identified, isolated and processed on separate layers.
In SpectraLayers Pro 12, I can separate a finished recording into vocals, drums, bass, guitar, piano, saxophone, brass and other material. Its Unmix Multiple Voices function can also distinguish different voices, including overlapping singers, by registering short voice profiles. Those vocal layers can then be processed independently and recombined afterwards.
I am not suggesting that VirtualDJ should reproduce the entire SpectraLayers environment in real time. SpectraLayers is effectively the Photoshop of audio editing. It provides precise surgical tools rather than the equivalent of a quick repair in the home garage.
The point is simply that modern audio software can already analyse complex, mixed material, distinguish multiple voices and process those voices separately. Multiple singers are therefore not an absolute technical barrier, although performing that degree of separation live and at sufficiently low latency would clearly be a different engineering challenge. It might also require more extensive pre-analysis of stems, similar to VirtualDJ’s existing optional stem preparation.
More importantly, VirtualDJ would not necessarily have to identify and isolate every singer individually.
A practical first implementation could apply formant-preserving pitch shifting to the combined vocal stem. The entire vocal stem would still be transposed by exactly the same musical interval as the instrumental stems. The additional processing would simply attempt to preserve the vocal stem’s spectral character while its pitch is changed.
Separating individual singers might improve the result in particularly difficult recordings, but it is not essential to the basic feature being proposed.
Regarding the amount of processing, I agree that tempo adjustment, keylock, key shifting, stem separation and formant preservation could create a demanding processing chain. However, these stages would not necessarily need to be implemented as four completely unrelated passes over the audio. A purpose-built algorithm could potentially combine parts of the pitch, time and formant processing within one integrated stage.
The result would not need to be flawless on every recording to be useful. An optional higher-quality mode that produces a more natural result on a substantial proportion of vocal material might still be a worthwhile improvement.
However, that creative use is slightly different from what I am proposing. Sometimes the chipmunk or monster effect is desirable. At other times, particularly when changing a track’s key for harmonic mixing, the aim is to retain as much of the original vocal character as possible. That is why I agree it should be optional.
The Roland VT-4 is a dedicated hardware unit with onboard DSP. Its architecture is designed to perform a limited set of pitch and formant processes in real time with very low latency. A contemporary laptop or MacBook has considerably greater overall processing capability, although of course it is also being asked to perform many more tasks at once.
VirtualDJ already separates stems and routes selected stems through independent effect-processing paths. It is therefore reasonable to ask whether a formant-preserving stage could eventually be integrated into the vocal stem’s existing processing path.
I currently work with SpectraLayers Pro 12. I have not upgraded to version 13 yet, as Steinberg often offers upgrade deals several months after release. Other software may be preferred by different users, of course.
Different strokes and all that jazz.
SpectraLayers is a professional spectral audio editor and source-separation application. Rather than presenting audio solely as a conventional waveform, it displays sound across time, frequency and amplitude. This allows individual elements within a recording to be identified, isolated and processed on separate layers.
In SpectraLayers Pro 12, I can separate a finished recording into vocals, drums, bass, guitar, piano, saxophone, brass and other material. Its Unmix Multiple Voices function can also distinguish different voices, including overlapping singers, by registering short voice profiles. Those vocal layers can then be processed independently and recombined afterwards.
I am not suggesting that VirtualDJ should reproduce the entire SpectraLayers environment in real time. SpectraLayers is effectively the Photoshop of audio editing. It provides precise surgical tools rather than the equivalent of a quick repair in the home garage.
The point is simply that modern audio software can already analyse complex, mixed material, distinguish multiple voices and process those voices separately. Multiple singers are therefore not an absolute technical barrier, although performing that degree of separation live and at sufficiently low latency would clearly be a different engineering challenge. It might also require more extensive pre-analysis of stems, similar to VirtualDJ’s existing optional stem preparation.
More importantly, VirtualDJ would not necessarily have to identify and isolate every singer individually.
A practical first implementation could apply formant-preserving pitch shifting to the combined vocal stem. The entire vocal stem would still be transposed by exactly the same musical interval as the instrumental stems. The additional processing would simply attempt to preserve the vocal stem’s spectral character while its pitch is changed.
Separating individual singers might improve the result in particularly difficult recordings, but it is not essential to the basic feature being proposed.
Regarding the amount of processing, I agree that tempo adjustment, keylock, key shifting, stem separation and formant preservation could create a demanding processing chain. However, these stages would not necessarily need to be implemented as four completely unrelated passes over the audio. A purpose-built algorithm could potentially combine parts of the pitch, time and formant processing within one integrated stage.
The result would not need to be flawless on every recording to be useful. An optional higher-quality mode that produces a more natural result on a substantial proportion of vocal material might still be a worthwhile improvement.
geposted Sun 26 Jul 26 @ 4:23 pm





