Spatial Audio & Immersive Music for Tuvan Throat Singing: Khoomei, Sygyt & Kargyraa
Spatial9 Team · · 12 min read
Part of the series Spatial audio and immersive music by genre.
Key takeaways
- In an immersive Tuvan throat singing mix the singer and the low drone sit close in the centre, the overtone melody floats just above and in front, and the igil, doshpuluur and khomus sit calmly to the sides.
- Because the brain separates sounds into streams by pitch and direction, a process called auditory stream segregation, giving the overtones a slightly different place from the drone helps listeners hear two lines from one voice.
- Keep every element still and use a wide, dark, open space with few early reflections, so the long notes can fade into distance without blurring the whistling harmonics.
- The Tuvan throat singing immersion profile scores width 3, height 4, movement 1, low-end focus 3 and space 4 out of 5.
Tuvan throat singing, called khöömei in Tuvan, is a way of singing in which one person sounds two or more pitches at once: a steady low drone and, above it, a melody made from the natural harmonics of the voice. It comes from Tuva, a republic of the Russian Federation in southern Siberia, on the northern border of Mongolia and around the upper Yenisei River. For generations Tuvans lived as herders, moving with their animals between steppe, forest and mountain pastures, and throat singing grew from this life. Singers imitate the sounds of the land, such as wind, flowing water and the calls of animals, and sing to the landscape as much as to other people. Similar traditions exist in western Mongolia and the Altai, but Tuva developed a rich set of named styles, among them khoomei, sygyt, kargyraa, borbangnadyr and ezengileer. The voice is often joined by the igil, a two-stringed bowed fiddle, the doshpuluur, a long-necked plucked lute, the byzaanchy, a bowed fiddle with four strings, the khomus, or jaw harp, and the dünggür, the frame drum of the shaman. In 1987 the American ethnomusicologist Theodore Levin and colleagues recorded singers in Tuva for the album Tuva: Voices from the Center of Asia, released by Smithsonian Folkways in 1990. The group Huun-Huur-Tu formed in Kyzyl in 1992 and toured the world, and the documentary Genghis Blues (1999) followed the American blues singer Paul Pena to Tuva, where he sang with the master Kongar-ol Ondar. Throat singing suits spatial audio because the music is about one voice in a vast landscape, and the overtones seem to float apart from the singer. This guide covers how spatial audio for Tuvan throat singing works, where to place each element and how to make an immersive Tuvan throat singing mix. For the basics, see what spatial audio is.
Why Tuvan throat singing works in spatial audio
Tuvan throat singing is built on very few parts. A drone holds steady in the throat while the singer shapes the mouth, tongue and lips to pick out one harmonic at a time, so a flute-like melody appears above the drone. Instruments add a gentle, often horse-like rhythm, and a second singer may join or answer. In stereo the drone and the overtone melody come from the same point, and many listeners hear only a strange, buzzing voice. In an immersive mix the singer stays centred and still, while the overtones sit a little higher and in front, so the ear can hear the two lines as separate voices. A wide, quiet space around the music gives the long notes somewhere to travel.
The rule: one still voice, with its overtones floating above it. Keep the singer and the drone centred and close, lift the overtone melody gently above ear level and set the small group of instruments inside a wide, open space that does not move.
Immersive music for Tuvan throat singing: what listeners hear
- Two voices from one singer. The drone sits low in the centre and the overtone melody floats slightly above it, so the listener hears both lines at once.
- A close, physical voice. The singer is placed near the listener, so every rasp, breath and change in the mouth is clear.
- A few instruments around the voice. The igil, doshpuluur and khomus sit to the sides, adding rhythm and colour without crowding the singer.
- Stillness and distance. Nothing moves around the listener. The space opens wide and far, like a valley or a stretch of steppe, and the long notes fade slowly into it.
Where to place each element in a Tuvan throat singing mix
| Element | Position | Height | Why it works |
|---|---|---|---|
| Throat singer | Front centre | Ear level | The voice that carries both drone and melody |
| Fundamental drone | Centre | Low | The steady low pitch under every phrase |
| Overtone melody | Front right | Raised | Whistling harmonics that seem to float above |
| Igil | Front left | Ear level | Bowed fiddle that doubles or answers the voice |
| Doshpuluur | Right | Ear level | Plucked lute that gives the riding rhythm |
| Khomus | Left | Ear level | Jaw harp with its own buzzing overtones |
| Dünggür drum | Behind left | Ear level | A soft frame drum pulse |
| Second voice | Behind right | Ear level | A partner singer joining or answering |
| Open steppe | Behind | Raised | A wide, quiet space with distant echoes |
The psychoacoustics behind Tuvan throat singing placement
Every position in the table above follows how the ear and brain locate sound. Here is the reason behind the main choices for Tuvan throat singing.
- The drone stays centred. The ear finds almost no direction in frequencies below roughly 80 to 100 Hz, so the lowest part of a deep kargyraa drone gains nothing from being spread. Keeping it centred makes the voice feel grounded and steady (Blauert).
- Overtones heard as a separate line. The brain groups sounds into streams by pitch, timing and direction. This is auditory stream segregation (Bregman). Giving the overtone melody a slightly different position from the drone helps the listener hear it as its own voice rather than as a buzz in the tone.
- High overtones feel higher. The ear judges height mostly from how the outer ear shapes high frequencies. These are spectral cues for elevation (Hebrank and Wright). Bright, whistling sygyt harmonics already carry some of this lift, so a small raise is enough to make them float.
- A wide space with side reflections. Early reflections arriving from the sides make a sound feel wider and more enveloping. This is spatial impression from early lateral reflections (Barron and Marshall). A few soft, distant reflections give the voice the sense of an open landscape without blurring the overtones.
The immersion profile for Tuvan throat singing
Width scores 3 because the music centres on one or two singers with a few instruments, so it does not need a full circle of sound. Height scores 4, since the high overtone melody naturally seems to rise above the voice, and lifting it makes the effect clear. Low-end focus is 3: the drone is deep, and kargyraa can reach very low, but the music has no heavy bass instruments and its weight sits in the voice. Space is 4, because the songs come from open land, with rivers, valleys and hills, and a wide sense of distance is part of the sound. Movement is 1, as the singer stands still and the music is calm, so sounds should stay fixed in place.
Tips for an immersive Tuvan throat singing mix
- Keep the singer close and centred. Throat singing is intimate and physical, so let the listener hear the breath and the texture of the drone.
- Lift the overtones a little. If you can isolate the overtone melody, raise it slightly above and in front of the voice. A small lift works better than a large one.
- Keep everything still. Do not pan or move the voice. The calm of the music depends on a fixed, steady image.
- Give the instruments their own sides. Put the igil and khomus on one side and the doshpuluur on the other, so their own overtones do not mask the singer's.
- Use a wide, dark space. A long, soft reverb with few early reflections suits the music, but keep it low in level and avoid bright tails that compete with the whistling harmonics.
Tuvan throat singing styles: Khoomei, Sygyt and Kargyraa
The core placement holds for every style here: the singer and drone at the centre, the overtones above and a wide, still space around the listener. What changes is the pitch of the drone, how focused the overtones are and how much low end the style carries.
Khoomei
Khoomei is both the general Tuvan word for throat singing and the name of one particular style. As a style it has a mid-range drone and soft, breathy overtones that sound diffuse, sometimes compared to wind swirling among rocks. It is often the first style a singer learns, and singers may move between khoomei and other styles within a single song. Because the overtones are gentle and spread across several harmonics, place the singer close and centred and use only a slight raise for the upper layer. A warm, open space with a little distance suits the style, with the igil to the front left answering the voice.
Sygyt
Sygyt, sometimes spelled sigit, means whistling, and it is the style with the clearest overtone melody. The singer holds a mid-range drone and presses the tongue toward the roof of the mouth to focus a single, very high harmonic, which sounds like a flute or a whistle above the voice. Sygyt melodies often evoke birdsong and summer breezes, and the effect is strongest when the overtone is easy to follow. Keep the drone centred and close, and lift the overtone melody a little higher and to the front, where the ear can track it without effort. Avoid bright reverb, which can blur the whistle. A clean, wide space with soft distant reflections suits it best.
Kargyraa
Kargyraa is the deepest style. The singer vibrates the ventricular folds, sometimes called the false vocal folds, together with the vocal folds, producing a growling tone that can sound about an octave below the normal voice. By changing the shape of the mouth, often through open vowels, the singer brings out a melody of low and mid harmonics. Tuvan singers speak of different kinds of kargyraa, such as steppe kargyraa and mountain kargyraa. Paul Pena taught himself kargyraa from recordings before travelling to Tuva. Keep the deep drone low and centred, close to the listener, with the overtones at ear level or only just above. Kargyraa carries more low end than the other styles, so keep the reverb short in the low frequencies.
Related styles such as borbangnadyr, with its rolling, trilling sound, and ezengileer, named after the stirrup and its horse-riding rhythm, follow the same idea, as do the overtone traditions of western Mongolia and the Altai: keep the voice centred, let the overtones float and give the music an open landscape.
How to make immersive Tuvan throat singing from stereo
Spatial9 can create spatial audio from a stereo track in a few steps:
- Upload your stereo master or your stems.
- Spatial9's AI separates and places each part, pulling the voice, igil, doshpuluur, khomus and drum apart even from field recordings, and keeping the singer centred with the instruments set calmly around.
- Preview it in binaural on headphones.
- Export Eclipsa Audio, ADM, binaural or Ambisonics.
- Publish a video. Create a YouTube-ready video with your immersive audio and publish it directly.
For related guides, read spatial audio for choral music and spatial audio for ambient music. Every guide is listed in spatial audio and immersive music by genre.
Frequently asked questions
Does this guide cover Khoomei, Sygyt and Kargyraa?
Yes. Khoomei keeps a close, centred voice with soft overtones only slightly raised. Sygyt lifts its whistling overtone melody a little higher and to the front. Kargyraa keeps its deep drone low and centred, with a shorter low end in the reverb.
Can I make a Tuvan throat singing mix in Dolby Atmos or Eclipsa Audio?
Yes. The placements in this guide work for both. Spatial9 exports ADM, the standard file used by Dolby Atmos studios, as well as Eclipsa Audio for YouTube, so one immersive Tuvan throat singing mix covers both. If you plan to release the Atmos version on Apple Music, build it from your original stems or multitracks, because Apple does not accept Atmos made from a finished stereo master. See where to publish spatial audio. You may also hear this called a 3D throat singing mix or khoomei in spatial audio.
Is spatial audio good for Tuvan throat singing?
Yes. Throat singing is one voice making two lines, set in a wide landscape. Spatial audio can give the overtone melody its own place above the drone and surround the voice with open space, which makes both lines easier to hear.
Can spatial audio really separate the overtones from the drone?
Partly. The overtones are part of the same voice, so they cannot always be fully pulled apart. When they can, a small lift helps. When they cannot, a close, centred voice in an open space still makes the overtones stand out.
Should I add movement to make it more immersive?
No. Tuvan throat singing is calm and steady, and moving the voice would distract from the overtones. Keep every element still and let the space and the melody provide the sense of depth.
How is Tuvan throat singing different from choral music in spatial audio?
Tuvan throat singing is usually one or two voices with a few instruments, and its interest lies in the overtones of a single voice, so it suits a close, centred singer in an open landscape. Choral music has many voices spread across sections in a resonant hall, so it suits a wider front arc and a longer, richer reverb.
Can I convert an old Tuvan field recording to spatial audio?
Yes. Spatial9 uses AI to separate the voice, igil, doshpuluur and khomus from a stereo or mono recording. Separation is never perfect with outdoor recordings and wind noise, so keep the layout simple and the space moderate.
Dolby and Dolby Atmos are trademarks of Dolby Laboratories. Spatial9 is not affiliated with or endorsed by Dolby Laboratories.
Sources
- Blauert, J. (1997). Spatial Hearing: The Psychophysics of Human Sound Localization, revised edition. MIT Press. Low-frequency, elevation, distance and motion cues.
- Mills, A. W. (1958). On the minimum audible angle. Journal of the Acoustical Society of America, 30(4), 237-246.
- Hebrank, J. and Wright, D. (1974). Spectral cues used in the localization of sound sources on the median plane. Journal of the Acoustical Society of America, 56(6), 1829-1834.
- Bronkhorst, A. W. (2000). The cocktail party phenomenon: a review of research on speech intelligibility in multiple-talker conditions. Acta Acustica united with Acustica, 86, 117-128.
- Barron, M. and Marshall, A. H. (1981). Spatial impression due to early lateral reflections in concert halls. Journal of Sound and Vibration, 77(2), 211-232.
- Neuhoff, J. G. (1998). Perceptual bias for rising tones. Nature, 395, 123-124.
- Bregman, A. S. (1990). Auditory Scene Analysis: The Perceptual Organization of Sound. MIT Press.
- Levin, T. with Süzükei, V. (2006). Where Rivers and Mountains Sing: Sound, Music, and Nomadism in Tuva and Beyond. Indiana University Press.
Sing to the steppe
Upload a track to Spatial9 and hear the drone at the centre with the overtones floating above you in minutes. Then export it or publish it straight to YouTube.