I got frustrated with every webcam air-mouse project I could find on GitHub. They all have the same fundamental usability problem: you hover your cursor over a button, you tap your fingers together to click, and the cursor ends up 60–80 pixels below the button you were aiming at. It's not a calibration issue. It's not a sensitivity issue. It's a timing problem baked into how click detection fundamentally works.
I built Gesture Mouse Controller specifically to fix this, and the solution took me a while to see.
Why every air-click misses
Here's what happens when you use a flick gesture (rapid downward finger motion) to trigger a click:
- Your hand is hovering over the target. The cursor is positioned correctly.
- You initiate a downward flick with your index finger.
- Your
FlickDetectoris computing average downward velocity across recent frames:vel = y_norm - self.prev_y. - After 3–5 frames of acceleration,
avg_velfinally crosses the thresholdFLICK_VEL_THRESHOLD = 0.022. - Click detected!
pyautogui.click(current_x, current_y)is called.
The problem is step 4 to 5. The instant detection fires, your finger has already traveled 5–8% of the camera frame downward from where it started the flick. Because the cursor tracking loop is running every frame with pyautogui.moveTo(int(smooth_x), int(smooth_y)), the cursor has been faithfully following your finger the entire time. By the time click fires, you're already 80 pixels south.
The cursor is doing exactly what it's supposed to. That's the problem.
The onset snapshot fix
The solution is to separate the trigger from the target location. Instead of clicking wherever the cursor is when detection fires, you snapshot the cursor position at the very onset of the flick motion — the first frame where velocity begins to climb — and use that position for the actual click dispatch.
if avg_vel > FLICK_VEL_THRESHOLD and not self.in_flick:
if (now - self.last_flick_t) > FLICK_COOLDOWN:
self.in_flick = True
self.last_flick_t = now
if self.pending_click and (now - self.pending_t) < DOUBLE_CLICK_WINDOW:
# Second flick within window — double-click at FIRST snapshot
action = 'double'
snap = self.pending_snapshot
self.pending_click = False
else:
# First flick — snapshot position RIGHT NOW before finger moves
self.pending_click = True
self.pending_t = now
self.pending_snapshot = screen_xy # <- the key line
# Flush pending single click once double-click window expires
if self.pending_click and (now - self.pending_t) >= DOUBLE_CLICK_WINDOW:
action = 'click'
snap = self.pending_snapshot # dispatch at saved position
self.pending_click = False
self.pending_snapshot = screen_xy at flick onset. That's the whole fix. When the click eventually dispatches, pyautogui.click(*snap) uses the pre-flick coordinates. You aim, you flick, the click lands exactly where you were pointing before your finger moved.
The 0.45-second DOUBLE_CLICK_WINDOW also handles double-clicks correctly: the second flick within the window dispatches a double-click at the first flick's snapshot position, not the current cursor position. This means you can double-click a file icon without the cursor drifting between taps.
Pinch-to-drag with hysteresis
Dragging was a separate headache. The basic approach: compute Euclidean distance between thumb tip (landmark 4) and index tip (landmark 8), normalized by palm scale (dist(lm[0], lm[5])). Below PINCH_THRESHOLD = 0.052 for 300ms → mouseDown. Above release threshold → mouseUp.
The problem without hysteresis: when you're dragging a window across the screen, your hand shifts and finger distance naturally fluctuates. A naive symmetric threshold drops the drag every time your fingers drift slightly apart mid-motion.
Hysteresis fixes this with two different thresholds. Entry: pinch below 0.052. Exit: pinch must open past PINCH_RELEASE = 0.075 before drag releases. The gap between those two numbers is a dead zone that absorbs natural hand instability during sustained drag. You can drag files across dual monitors without worrying about accidental drops.
Cursor smoothing and the lerp factor
Raw index fingertip coordinates from MediaPipe jitter constantly — hand tremor, webcam sensor noise, occasional landmark misdetections. Piping raw coordinates directly into pyautogui.moveTo() produces a cursor that vibrates visibly even when you're holding your hand perfectly still.
The fix is exponential moving average lerp:
smooth_x = lerp(smooth_x, nx * screen_w, 0.18) smooth_y = lerp(smooth_y, ny * screen_h, 0.18)
A lerp factor of 0.18 was the sweet spot after testing values from 0.05 (too laggy, cursor floats to catch up) to 0.5 (too twitchy, barely better than raw). At 0.18, you can flick across a dual-monitor setup quickly, and you can also hold a 16-pixel desktop icon steady enough to click it.
The thing that's still not fixed
Rapid saccadic hand movements — where you snap your hand to a new position instantly — produce a lerp lag "tail" where the cursor takes two or three frames to catch up. It's usually fine, but it's noticeable when you're moving between far-apart targets quickly. A velocity-adaptive lerp factor (higher factor when velocity is high, lower when velocity is low) would fix this, but I haven't implemented it yet.
Takeaway
In gesture recognition, your biological action takes time to reach classification confidence. If you apply the result of a gesture to the instantaneous state of the hand when detection succeeds, you'll always be late and off-target. Record the pre-gesture snapshot at onset and execute your action retroactively. The gesture tells you what to do; the onset position tells you where to do it.