While building my ESP32 ultrasonic obstacle detector, I realized a fundamental hardware limitation: ultrasound gives you a single one-dimensional distance reading along a narrow cone. It tells you that something is 40 centimeters in front of you, but it can't tell you whether that object is a low concrete curb, an overhanging tree branch, a glass door, or a moving dog. Drishti (the Sanskrit word for "vision") was my attempt to bridge that gap by building a real-time computer vision and spatial awareness engine that runs on portable consumer hardware.
The primary engineering obstacle in assistive computer vision is the brutal trade-off between semantic accuracy and pipeline latency. If a person is walking at a standard pace of 1.4 meters per second, an obstacle detection delay of even 300 milliseconds means they've taken half a step into a potential hazard before the alert sounds. When I first chained an object detection network to a monocular depth estimation model in Python, the pipeline choked at 8 frames per second, setting my laptop fans screaming and lagging audio feedback by nearly half a second.
To make the system viable for real-world movement, I completely decoupled the architecture into two asynchronous, multi-threaded pipelines:
- High-Frequency Spatial Tracking (60 Hz): A dedicated lightweight vision thread runs sparse optical flow (Lucas-Kanade) and frame-differencing bounding-box tracking. It monitors high-velocity motion vectors and impending collision trajectories with near-zero latency.
- Low-Frequency Semantic Classification (10 Hz): A background worker thread processes downscaled frames through an int8-quantized YOLO model, labeling obstacle classes (vehicles, pedestrians, stairs, doors, street signs). Bounding box classifications are asynchronously projected onto the active spatial tracker.
The next major hurdle was floor false positives. Because the camera is angled forward, standard depth algorithms constantly flag the ground 2 meters ahead as an oncoming obstacle. To fix this, I implemented a RANSAC (Random Sample Consensus) ground-plane fitting algorithm in OpenCV. It models the dominant planar surface beneath the user's eye level, calculates the floor normal vector, and mathematically strips out points lying on the walking plane. This ensures that ground texture, shadows, and carpet seams are ignored while actual drop-offs, potholes, and elevated curbs are immediately flagged.
Finally, visual telemetry is translated into human intuition through a 3D binaural spatial audio engine. Using stereo sound synthesis, obstacle coordinates (X, Y, Z) are mapped directly into headphones: an obstacle on the left pans hard left, elevation shifts the acoustic harmonic timbre, and physical proximity increases both pitch and pulse repetition frequency. It lets the user perceive the physical geometry of a room entirely through sound.