Skip to the record

CASE: Voice-Controlled Self-Balancing Robot

2026 · ESP32 · FreeRTOS · C/C++ · PID control · MPU6050 DMP · I²S audio

Design and build a two-wheeled self-balancing robot that holds vertical equilibrium via real-time PID control on a dual-core ESP32, integrating on-device wake-word voice input with cloud-based conversational AI so the balance loop is never interrupted by audio processing.

100 Hz
Balance Loop Rate
2
Dedicated Cores

First look


Render of the CASE robot: a rounded upright body on two wheels, with a level-meter display on its face
Render of the CASE robot: a rounded upright body on two wheels, with a level-meter display on its face

What it is made of


Hardware

  • Ran the balance loop on an ESP32-WROOM-32's Core 1 under FreeRTOS, so WiFi and TLS never share a core with it. The loop holds 100 Hz with under 10 µs of jitter.
  • Read attitude from an MPU6050 through its on-chip DMP quaternion fusion, so the controller receives a drift-corrected angle without spending MCU cycles on filtering.
  • Drove two DC motors through a TB6612FNG off a 7.4 V pack, so PID output becomes PWM with the current headroom the chassis needs to catch itself.
  • Captured voice on an INMP441 I2S microphone and answered through a MAX98357A into a 4 Ω 3 W speaker, so the robot hears and replies without an external codec.

Software

  • Held vertical equilibrium with a PID controller (Kp 15 · Ki 140 · Kd 0.9) fed by the DMP at 100 Hz, so the robot stays upright while it is talking.
  • Detected “Hey CASE” on-device from the I2S stream, so the network is only touched once someone has actually spoken.
  • Bridged audio to Nvidia Personaplex (Moshi) through a Google Colab endpoint transcoding PCM ↔ Opus, so conversational AI runs off-board inside a 167 KB heap.

Why the loop stays on the microcontroller


The same gains, the same shove, two schedulers. On the left the control update runs when the timer fires; on the right it runs when a general-purpose operating system gets round to it: mostly a few milliseconds late, occasionally 140 ms late. Watch the right-hand chassis wander.

ESP32 · interrupt-driven loopruns when the timer fires · jitter under 10 µs
outside the band
0.00 s
now
4.98°
rests at
0.16°
The same code under a general-purpose schedulermostly 2–15 ms late · occasionally the CPU is elsewhere for 140 ms
outside the band
0.00 s
now
4.98°
rests at
0.32°

Lean drawn ×6. The difference between these two is fractions of a degree, which is the whole point and invisible at life size.

Neither falls, which is the honest result. The reason to keep the loop on the microcontroller is margin, not survival. Under the scheduler the same gains spend roughly three times as long outside the half-degree band and come to rest further off upright, and every one of those excursions is time the chassis is spending its recovery budget on the operating system instead of on the next shove.

What came of it


  • The robot holds vertical equilibrium while handling live voice input.
  • The balance loop is never interrupted by audio or network work.

The trade it forced


Why the ESP32 alone wasn’t enough

167 KB of free heap after WiFi and TLS caps the board at 1.5 s recordings and HTTP round-trips, with no streaming anywhere. But the interrupt-driven DMP loop holds <10 µs jitter where a Linux scheduler gives 2–15 ms, so balance stays on the MCU and only the voice pipeline moves to a Raspberry Pi. About 80% of self_balance.ino survives that migration.

The other plates


KiCad layout of the robot’s board: the ESP32-WROOM-32 against its keep-out zone, the TB6612FNG and buck regulator across the middle, and the I2C, I2S and UART headers along the bottom edge
3D render of the fabricated board: the ESP32-WROOM-32 module, electrolytic capacitors, the TB6612FNG driver, green screw terminals for both motors, and the labelled header rows

Self-balancing · Wake-word activation · Conversational AI · Voice-commanded motion · Audio capture & playback