Running computer vision models on a high-powered desktop with an active GPU is relatively easy. Running real-time machine learning inference directly inside the constrained SRAM of a low-power microcontroller without any internet connection, cloud API, or external camera is a completely different engineering challenge. I built the TinyML Gesture Circle Detector to explore edge machine learning by classifying physical 3D spatial hand gestures using an MPU6050 6-axis Inertial Measurement Unit (IMU).
The system interfaces with the MPU6050 over I2C, reading 3-axis accelerometer data (X, Y, Z linear acceleration) and 3-axis gyroscope data (roll, pitch, yaw angular velocity). The on-device inference pipeline operates in three tight, deterministic stages:
- Hardware Timer & Sliding Window Buffer: Gesture classification requires temporal context. A hardware timer interrupt polls the IMU at a locked 25 Hz sampling rate, writing sensor readings into a circular sliding window buffer holding 50 consecutive frames across all 6 motion axes. This forms a flattened 300-dimensional feature vector (
50 samples × 6 axes) representing the trailing two seconds of physical motion. - On-Chip Feature Standardization: Raw IMU readings carry gravitational offsets, sensor bias, and physical drift. Before feeding data to the neural network, the firmware standardizes the 300 features against pre-calibrated floating-point mean (
feature_mean) and standard deviation (feature_std) lookup tables stored in flash memory. - Quantized Neural Network Inference: The model is executed via the ArduTFLite (TensorFlow Lite for Microcontrollers) runtime. The network architecture uses a compact dense classifier quantized from 32-bit floating point down to 8-bit integers (
int8), stored as a constant byte array inmodel_data.h.
The most brutal debugging hurdle of this build was memory alignment. In my early C++ sketches, the microcontroller would randomly crash with an undocumented LoadProhibited hardware panic during tflite::MicroInterpreter::AllocateTensors(). After hours of digging through low-level assembly dumps and memory maps, I discovered that TensorFlow Lite Micro strictly requires 16-byte memory boundary alignment for SIMD vector operations. Allocating a standard uint8_t tensorArena[16 * 1024] in ESP32 SRAM caused memory misalignment faults. The fix was forcing memory alignment at compile time:
alignas(16) uint8_t tensorArena[16 * 1024];
Once the arena was properly aligned, the quantized model executed in under 12 milliseconds per inference pass, reliably recognizing deliberate circular clockwise and counter-clockwise hand gestures while discarding ambient arm jitter and walking vibrations.