BREAKING: IntentResolver now resolves intents by SEEING the screenshot
instead of parsing XML text descriptions.
Architecture:
- PRIMARY: Visual Discovery (SoM) — annotates screenshot with numbered
bounding boxes, sends to VLM, VLM visually picks the right box
- FALLBACK: Text-based VLM resolution (only when no device available)
- Removed: _visual_critic (redundant — visual discovery IS visual)
- Removed: _humanize_desc regex (the VLM reads the actual screen now)
Key innovations:
- Spatial Deduplication: child nodes fully contained in parent bounds
are suppressed (83 → ~19 boxes), eliminating visual noise
- System UI filtering: statusbar, notifications excluded from candidates
- VLM prompt is pure visual: 'look at the numbered boxes and pick one'
Proven by live LLM test: VLM correctly identifies 'following' (not
'followers') by SEEING the screen content, with zero string matching.