The setup
The final project for CS 6476, Computer Vision, at Georgia Tech (Spring 2026): classify walking, jogging, running, boxing, handwaving and handclapping in the KTH dataset (599 clips, 25 people). No neural networks, no pretrained models, and no cv2.HuMoments, so every invariant is written out from its definition. I chose explicit features on purpose, so that when the model is wrong I can see which feature pushed it there.
What I built
Frame differencing into motion masks (close, then open), motion history images with a per-action memory length, and Hu invariants computed across eight threshold and kernel channels stacked into one feature row, so the random forest learns which channel matters for which pair of actions. Tuned with leave-one-subject-out cross-validation, tested on nine held-out subjects.
The part I would do again: the 56-feature baseline hit 70%, and three different ensemble families tied within 2.4 points of it, which said the ceiling was the features, not the classifier. Per-class SHAP pointed at three specific gaps, and each gap got one feature family: scalar shape and mass (boxing went from 0.58 to 0.81), half-silhouette moments, and an FFT of motion area over time for cadence.
Numbers
| Features | Test accuracy |
|---|---|
| 56 (Hu only) | 0.7037 |
| 88 (+ shape and mass) | 0.8333 |
| 200 (+ half-silhouette) | 0.8519 |
| 240 (+ periodicity) | 0.8611 |
Mean across eight seeds 0.8640 (sd 0.0078). None of the most confidently wrong clips goes above 0.65 confidence: where it fails, it fails unsure, at the boundaries where the actions themselves are close.
Course policy keeps the code private, so the write-up is the whole public artifact.