# FMOD Voice Recording access to real-time playable bytes/data (C++)

**URL:** <https://qa.fmod.com/t/fmod-voice-recording-access-to-real-time-playable-bytes-data-c/22311>\
**Category:** FMOD Engine\
**Created:** [November 15, 2024, 5:52pm UTC](https://qa.fmod.com/t/fmod-voice-recording-access-to-real-time-playable-bytes-data-c/22311 "2024-11-15T17:52:51Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![CreativePS](https://avatars.discourse-cdn.com/v4/letter/c/c89c15/32.png) [@CreativePS](https://qa.fmod.com/u/CreativePS)\
**Post date:** [November 15, 2024, 5:52pm UTC](https://qa.fmod.com/t/fmod-voice-recording-access-to-real-time-playable-bytes-data-c/22311/1 "2024-11-15T17:52:51Z")

</div>

Hello,

I’ve been using FMOD recently and have already looked at the record.cpp example.

However, in order to capture the audio data, it advises to use FMOD:🔉:lock and FMOD:🔉:unlock in sort of an external way, even though recordStart is non-blocking.

How exactly would I fetch the bytes in a synchronized way, so that there is no popping or artifacts? I’m aware that FMOD is internally using circular buffers as well. Is there no callback I could use? I’ve tried FMOD::System::CreateStream and tried to point recordStart to there and just use PCMReadCallback, which then complains about:

“[ERR] SystemI::recordStart : Invalid sound, must be an FMOD::Sound with positive length created as FMOD\_CREATESAMPLE.”, even when I do fulfill everything that FMOD complains about.

Could anyone help? Also, I’m using C++ to do this.

Thank you!

---

<div class="post-metadata">

**Author:** ![li\_fmod](https://yyz2.discourse-cdn.com/flex036/user_avatar/qa.fmod.com/li_fmod/32/6577_2.png) [@li\_fmod](https://qa.fmod.com/u/li_fmod)\
**Post date:** [November 20, 2024, 5:26am UTC](https://qa.fmod.com/t/fmod-voice-recording-access-to-real-time-playable-bytes-data-c/22311/2 "2024-11-20T05:26:48Z")

</div>

Hi,

Thank you for sharing the information.

> [@CreativePS](#):
>
> How exactly would I fetch the bytes in a synchronized way, so that there is no popping or artifacts?

It seems that `lock` and `unlock` are unable to access the audio buffer, you could consider using a custom DSP to capture the audio buffer directly and add it to your recording channel by using [ChannelControl::addDSP](https://www.fmod.com/docs/2.02/api/core-api-channelcontrol.html#channelcontrol_adddsp).

I will share an example script below as a reference:

```auto
FMOD_DSP_DESCRIPTION tap{};
tap.pluginsdkversion = FMOD_PLUGIN_SDK_VERSION;
tap.numinputbuffers = 1;    
tap.numoutputbuffers = 1;   
tap.numparameters = 0;  

tap.process = [](FMOD_DSP_STATE *dsp_state, unsigned int numsamples, const FMOD_DSP_BUFFER_ARRAY *inbufferarray,
                 FMOD_DSP_BUFFER_ARRAY *outbufferarray, FMOD_BOOL inputsidle, FMOD_DSP_PROCESS_OPERATION op) -> FMOD_RESULT {
    if (op == FMOD_DSP_PROCESS_PERFORM) {
        if (inputsidle) {
            // If input is idle, optionally clear the output buffer to silence
            memset(outbufferarray->buffers[0], 0, numsamples * sizeof(float));
            return FMOD_OK;
        }

        // Process audio samples
        for (unsigned int i = 0; i < numsamples; ++i) {
            outbufferarray->buffers[0][i] = inbufferarray->buffers[0][i] * 0.5f; // Reduce amplitude to prevent clipping
        }
    }

    return FMOD_OK;
};

    FMOD::DSP *myDSP;
    result = system->createDSP(&tap, &myDSP);
    ERRCHECK(result);

    result = myRecordingChannel->addDSP(0, myDSP); // Add DSP to the channel
    ERRCHECK(result);

```

> [@CreativePS](#):
>
> Is there no callback I could use? I’ve tried FMOD::System::CreateStream and tried to point recordStart to there and just use PCMReadCallback, which then complains about:
> 
> “[ERR] SystemI::recordStart : Invalid sound, must be an FMOD::Sound with positive length created as FMOD\_CREATESAMPLE.”, even when I do fulfill everything that FMOD complains about.

Unfortunately, FMOD does not support using `PCMReadCallback` for recording since [recordStart](https://fmod.com/docs/2.02/api/core-api-system.html#system_recordstart) requires a sound object created with `FMOD_CREATESAMPLE`, which is incompatible with [creatStream](https://fmod.com/docs/2.02/api/core-api-system.html#system_createstream) that is meant for streaming playback.

Hope this helps, let me know if you have any questions.

---

<div class="post-metadata">

**Author:** ![CreativePS](https://avatars.discourse-cdn.com/v4/letter/c/c89c15/32.png) [@CreativePS](https://qa.fmod.com/u/CreativePS)\
**Post date:** [November 20, 2024, 11:49am UTC](https://qa.fmod.com/t/fmod-voice-recording-access-to-real-time-playable-bytes-data-c/22311/3 "2024-11-20T11:49:56Z")

</div>

> [@li\_fmod](#):
>
> ```auto
> FMOD_DSP_DESCRIPTION tap{};
> tap.pluginsdkversion = FMOD_PLUGIN_SDK_VERSION;
> tap.numinputbuffers = 1;    
> tap.numoutputbuffers = 1;   
> tap.numparameters = 0;  
> 
> tap.process = [](FMOD_DSP_STATE *dsp_state, unsigned int numsamples, const FMOD_DSP_BUFFER_ARRAY *inbufferarray,
> FMOD_DSP_BUFFER_ARRAY *outbufferarray, FMOD_BOOL inputsidle, FMOD_DSP_PROCESS_OPERATION op) -> FMOD_RESULT {
> if (op == FMOD_DSP_PROCESS_PERFORM) {
> if (inputsidle) {
> // If input is idle, optionally clear the output buffer to silence
> memset(outbufferarray->buffers[0], 0, numsamples * sizeof(float));
> return FMOD_OK;
> }
> 
> // Process audio samples
> for (unsigned int i = 0; i < numsamples; ++i) {
> outbufferarray->buffers[0][i] = inbufferarray->buffers[0][i] * 0.5f; // Reduce amplitude to prevent clipping
> }
> }
> 
> ```

Hi, what I essentially want is to just capture a real-time buffer in a way that I can just use sendto() on it using a simple UDP socket and just fire-and-forget it so that the audio (my voice) can be replayed in real-time on the other side, allowing for a simple voice cha functionality with FMOD.

---

<div class="post-metadata">

**Author:** ![CreativePS](https://avatars.discourse-cdn.com/v4/letter/c/c89c15/32.png) [@CreativePS](https://qa.fmod.com/u/CreativePS)\
**Post date:** [November 20, 2024, 11:50am UTC](https://qa.fmod.com/t/fmod-voice-recording-access-to-real-time-playable-bytes-data-c/22311/4 "2024-11-20T11:50:26Z")

</div>

Oh, and thank you for your response sir.

---

<div class="post-metadata">

**Author:** ![li\_fmod](https://yyz2.discourse-cdn.com/flex036/user_avatar/qa.fmod.com/li_fmod/32/6577_2.png) [@li\_fmod](https://qa.fmod.com/u/li_fmod)\
**Post date:** [November 22, 2024, 12:09am UTC](https://qa.fmod.com/t/fmod-voice-recording-access-to-real-time-playable-bytes-data-c/22311/5 "2024-11-22T00:09:59Z")

</div>

Unfortunately, FMOD does not provide built-in networking support, so you’ll need to handle the transmission yourself.

You could consider packaging audio samples `inbufferarray->buffers[0][i]` into a custom structure like this:

```auto
    struct Packet
    {
        char data[32];
        int datalen;
    };

```

Then, use your own networking API (e.g., UDP sockets) to send the data over the network. On the receiver end, unpack the data and play it back using FMOD.

If you are looking for more information, there was a discussion relate to this topic that is worth reading through:

> [@Using FMOD with voice chat utilizing OnAudioFilterRead()](https://qa.fmod.com/t/using-fmod-with-voice-chat-utilizing-onaudiofilterread/17763):
>
> I’m trying to integrate Normcore multiplayer SDK into our project. They offer low-latency voice chat by utilizing OnAudioFilterRead(), playing a dummy clip, and then use an audio effect to inject voice data. I’d like to use the FMOD snapshots in our project with the voice chat system - is this possible given the above info? Or is OnAudioFilterRead() strictly for use with the Unity audio engine and therefore not accessible with FMOD? Any guidance would be greatly appreciated. Thank you!

---

<div class="post-metadata">

**Author:** ![FmodUserStd](https://avatars.discourse-cdn.com/v4/letter/f/77aa72/32.png) [@FmodUserStd](https://qa.fmod.com/u/FmodUserStd)\
**Post date:** [December 22, 2024, 11:13pm UTC](https://qa.fmod.com/t/fmod-voice-recording-access-to-real-time-playable-bytes-data-c/22311/6 "2024-12-22T23:13:18Z")

</div>

Could you exactly make an example of how to unpack the data and play it back using FMOD? Do I have to do it through the DSP again? Otherwise, I have everything else already working and good to go.

---

<div class="post-metadata">

**Author:** ![li\_fmod](https://yyz2.discourse-cdn.com/flex036/user_avatar/qa.fmod.com/li_fmod/32/6577_2.png) [@li\_fmod](https://qa.fmod.com/u/li_fmod)\
**Post date:** [December 29, 2024, 10:11pm UTC](https://qa.fmod.com/t/fmod-voice-recording-access-to-real-time-playable-bytes-data-c/22311/7 "2024-12-29T22:11:28Z")

</div>

Hi, sorry for the late response.

> [@FmodUserStd](#):
>
> Could you exactly make an example of how to unpack the data and play it back using FMOD? Do I have to do it through the DSP again?

I’m not fully familiar with the specific details of your implementation, so I can provide a high-level explanation based on general FMOD practices.

To play back real-time audio data captured via FMOD, you would typically handle it through a custom DSP. Here’s a high-level example of how you might set this up:

1. **Create a custom DSP** : Define your DSP to capture incoming audio data. Use the [`FMOD_DSP_DESCRIPTION`](https://www.fmod.com/docs/2.02/api/plugin-api-dsp.html#fmod_dsp_description) structure to describe your DSP and define a callback for processing audio data ([`FMOD_DSP_READ_CALLBACK`](https://www.fmod.com/docs/2.02/api/plugin-api-dsp.html#fmod_dsp_read_callback)).
2. **Register and add the DSP to your audio system** : After defining your DSP, register it with [System::registerDSP](https://www.fmod.com/docs/2.02/api/core-api-system.html#system_registerdsp) and then add it to your desired audio channel using [Channel::addDSP](https://www.fmod.com/docs/2.02/api/core-api-channelcontrol.html#channelcontrol_adddsp) .
3. **Data Handling in Callback** : In your DSP read callback, use the [lock](https://www.fmod.com/docs/2.02/api/core-api-sound.html#sound_lock) and [unlock](https://www.fmod.com/docs/2.02/api/core-api-sound.html#sound_unlock) functions to access the captured audio buffer. Here you process or modify the audio data as needed.
4. **Playback** : To play back the audio data, you can directly write it to an output channel within the DSP process callback.

---

<div class="post-metadata">

**Author:** ![CreativePS](https://avatars.discourse-cdn.com/v4/letter/c/c89c15/32.png) [@CreativePS](https://qa.fmod.com/u/CreativePS)\
**Post date:** [December 30, 2024, 7:47pm UTC](https://qa.fmod.com/t/fmod-voice-recording-access-to-real-time-playable-bytes-data-c/22311/8 "2024-12-30T19:47:31Z")

</div>

Hi,

does that mean I have to use the playSound callback? Cause what I am doing right now is that I’m using recordStart to pass it to an FMOD sound object, then I play the sound on the sound object and obtain the channel to attach it to the DSP. Then, in order to avoid the user hearing himself speak, I mute the output in the callback by clearing the out buffers.

It seems wrong. I don’t want to rely on the output device at all for this, this should work even if the user doesn’t have an output device. How could I do it properly?

I also want to use opus for encoding so I can send the samples through a server and whatnot, but first and foremost I want to solve exactly what I have stated above…

Thank you, much respect for all help!

---

<div class="post-metadata">

**Author:** ![li\_fmod](https://yyz2.discourse-cdn.com/flex036/user_avatar/qa.fmod.com/li_fmod/32/6577_2.png) [@li\_fmod](https://qa.fmod.com/u/li_fmod)\
**Post date:** [January 5, 2025, 11:47pm UTC](https://qa.fmod.com/t/fmod-voice-recording-access-to-real-time-playable-bytes-data-c/22311/9 "2025-01-05T23:47:39Z")

</div>

Hi,

Thank you for the detailed explanation.

> [@CreativePS](#):
>
> does that mean I have to use the playSound callback?

Could you clarify what you mean by ‘playSound callback’? Are you referring to a specific FMOD callback, such as [FMOD\_STUDIO\_EVENT\_CALLBACK\_SOUND\_PLAYED](https://fmod.com/docs/2.02/api/studio-api-eventinstance.html#fmod_studio_event_callback_type) or something else?

> [@CreativePS](#):
>
> Cause what I am doing right now is that I’m using recordStart to pass it to an FMOD sound object, then I play the sound on the sound object and obtain the channel to attach it to the DSP. Then, in order to avoid the user hearing himself speak, I mute the output in the callback by clearing the out buffers.

Your implementation sounds reasonable to me, as it works independently of the audio output device. You can verify this by setting FMOD’s output mode to `OUTPUTTYPE_NOSOUND`.

> [@CreativePS](#):
>
> It seems wrong. I don’t want to rely on the output device at all for this, this should work even if the user doesn’t have an output device. How could I do it properly?

Could you please elaborate on how you using the output device directly? It seems your current approach relies on the FMOD mixer, which is independent of the output device and should function regardless of the selected output type, such as `OUTPUTTYPE_NOSOUND` (no sound) or `OUTPUTTYPE_WAVWRITER` (output to file).

> [@CreativePS](#):
>
> I also want to use opus for encoding so I can send the samples through a server and whatnot, but first and foremost I want to solve exactly what I have stated above…

Please note that you’ll need to handle Opus encoding yourself, as FMOD provides access only to the raw PCM buffer.

---

<div class="post-metadata">

**Author:** ![CreativePS](https://avatars.discourse-cdn.com/v4/letter/c/c89c15/32.png) [@CreativePS](https://qa.fmod.com/u/CreativePS)\
**Post date:** [February 3, 2025, 11:55pm UTC](https://qa.fmod.com/t/fmod-voice-recording-access-to-real-time-playable-bytes-data-c/22311/10 "2025-02-03T23:55:43Z")

</div>

Hi, so I am now able to capture and playback PCM audio samples.

However, I want to ask, how would I deal with the jitter buffer? I have been trying to use an external thread for the jitter buffer and make sure that once it’s filled with at least 3 samples, that it starts popping off each one at a rate of exactly 20 ms.

However, I also have a problem when multiple users join the conversation…  
The sounds start to become robotic,and I am playing back every single sample with fmod. Am I supposed to use streams or something?

What could be going on?

---

<div class="post-metadata">

**Author:** ![Connor\_FMOD](https://yyz2.discourse-cdn.com/flex036/user_avatar/qa.fmod.com/connor_fmod/32/3685_2.png) [@Connor\_FMOD](https://qa.fmod.com/u/Connor_FMOD)\
**Post date:** [February 10, 2025, 6:23am UTC](https://qa.fmod.com/t/fmod-voice-recording-access-to-real-time-playable-bytes-data-c/22311/11 "2025-02-10T06:23:06Z")

</div>

Hi,

> [@CreativePS](#):
>
> However, I want to ask, how would I deal with the jitter buffer? I have been trying to use an external thread for the jitter buffer and make sure that once it’s filled with at least 3 samples, that it starts popping off each one at a rate of exactly 20 ms.

I would suggest looking over our video playback example: [Unity Integration | Scripting Examples Video Playback](https://fmod.com/docs/2.02/unity/examples-video-playback.html). It is in C#, but the idea of reducing and increasing the pitch to account for delay may help.

> [@CreativePS](#):
>
> However, I also have a problem when multiple users join the conversation…  
> The sounds start to become robotic,and I am playing back every single sample with fmod. Am I supposed to use streams or something?

Would it be possible to get a recording when the audio becomes robotic? Could it possible be linked to the networking of the conversation?

Would it be possible to get access to the code?
