- Updated QNavGraph.do to accept **kwargs and forward them to GOAP.
- Added explicit 'type and post comment' intent handler in GoalExecutor.
- Implemented smart VLM-driven Meta AI tone selection. If 'Supportive', 'Funny', or other Meta AI chips are visible on the comment keyboard, the bot uses the VLM to pick the best tone chip and lets Meta AI write the comment, falling back to manual typing if no chips are found.
- Fixed AttributeError where device.get_hierarchy() was called instead of dump_hierarchy()
- Replaced MagicMock with a real DummyDeviceForKeyboardTest in E2E tests to adhere to the strict mock ban policy.
- Verified TelepathicEngine actively calls device.back() when keyboard is open on non-typing intents.
- Added active Keyboard Dismiss logic to TelepathicEngine.find_best_node():
If the keyboard is open and the intent is NOT typing-related (e.g., 'tap post username'),
the bot automatically issues a back press to close the keyboard and re-fetches the UI state.
- Broadened structural fast-path for 'tap send post button' in IntentResolver to
also match 'direct_share_button' or 'send'/'share' desc, preventing VLM fallback
on POST_DETAIL screens (Bug #3).
1. ScreenIdentity: Restructured priority cascade to prevent Qdrant semantic
cache from overriding deterministic structural heuristics. Cached types
now resolve AFTER structural checks (Priority 3) instead of before.
Story/Reels hallucinations from both cache and VLM are rejected if
structural markers are absent.
2. bot_flow: Added error handling for story ring avatar tap and curiosity
HomeFeed navigation failures — prevents silent continuation into
undefined states.
3. GOAP: Extended is_navigation detection to include ScreenTopology
structural actions, ensuring HD Map routes are correctly classified.
4. action_memory: Tightened _parse_yes_no to prevent JSON fall-through
into ambiguous text matching. Structured responses with 'success'
boolean are now handled directly.
The SAE perceive() had no structural fast-path for Instagram-internal
modal overlays (surveys, rating prompts, interstitials). These modals
live inside com.instagram.android and were invisible to the foreign-app
detector. The LLM fallback frequently misclassified them as NORMAL,
causing the bot to get trapped in survey dialogs.
Added two O(1) structural detection layers:
1. Resource-ID markers: survey_overlay_container, interstitial_container,
mystery_interstitial, nux_overlay, rating_prompt, feedback_dialog
2. Dismiss-button heuristic: cross-validates dismiss text with overlay
container structure to prevent caption false positives
Also removed dead code: xml_dump.lower() no-op.
Fixes: test_perceive_instagram_survey_modal
Zero regressions: 464 passed, 6 pre-existing failures
- device_emulator.py: Replace MagicMock(watcher) with explicit WatcherStub
that implements exact u2 watcher chain (when/click/start). Any API drift
now causes AttributeError instead of silent acceptance.
- test_system_sae.py: Inject XML into emulator state instead of bypassing
dump_hierarchy with lambda
- test_behavior_ad_guard.py: Pass XML to fixture, remove lambda override
- test_behavior_scrape_profile.py: Pass XML to fixture, remove lambda override
Result: ZERO unittest.mock imports, ZERO MagicMock instances,
ZERO dump_hierarchy lambdas across entire E2E suite (50 files).
Production bug 2026-05-02 22:19:
Intent 'post author username text (exclude bottom tabs)' contained the
word 'tab' in its parenthetical qualifier. The old check:
is_tab_intent = 'tab' in intent_lower
matched this intent as a TAB navigation intent, activating the Tab
Height Guard which excluded ALL nodes with center_y < 85% of screen.
This nuked the actual clips_author_username node at y=2012, leaving
only 4 irrelevant items for the VLM → bot stuck in a loop unable to
find the author username.
Fix: Replace substring match with precise regex patterns that only
match real tab navigation intents:
- 'tap profile tab', 'tap home tab' (tap + word + tab)
- 'profile tab', 'explore tab' (word + tab)
- 'tab ...' at start of intent
TDD: Red test added first, reproducing exact production state.
81 perception tests + 172 total tests pass.
CONTRACT 6: extract_json() — LLM Response Sanitizer (9 tests)
Previously COMPLETELY untested. Every single LLM response in the entire
system passes through this function. Tests:
- Clean JSON, markdown fences (json+plain), thinking block purge
- Text prefix before JSON, truncated JSON fuzzy recovery
- Garbage/empty/None resilience
CONTRACT 7: _visual_discovery Pre-Filters (8 tests)
The guards INSIDE _visual_discovery() that run BEFORE the VLM ever
sees the candidates. If they fail, the VLM gets garbage and hallucinates.
Tests:
- Area filter: tiny nodes (<200px²) and huge containers (>400000px²)
- SystemUI filter: Android system elements, notifications, battery
- Strict Button Guard: long text captions excluded for button intents
- Grid Item Guard: non-grid elements excluded, fallthrough on no match
- Author/Username Guard: nav tabs excluded for author intents
CONTRACT 8: verify_success Structural Delta (8 tests)
The post-click verification logic that decides if a click actually worked.
Tests:
- Toggle massive XML shift = navigation error (False)
- Toggle small diff = success (True)
- Toggle zero diff = inconclusive (None)
- Follow success with 'Requested' marker (private accounts)
- Follow success with German locale ('Abonniert')
- View post with clips_viewer marker (Reels)
- View post still on grid = inconclusive (None)
- Semantic gate blocks wrong toggle element (caption != like button)
Two critical test additions:
1. test_system_full_lifecycle.py — Full A-to-Z integration test:
Exercises the ENTIRE bot pipeline across 4 screen types (HomeFeed →
Explore → Post Detail → Profile) with a real InstagramEmulator state
machine. Validates plugin execution, session state accumulation,
cognitive stack survival, dopamine session timeout, and post-session
serialization. This is the single most important test in the suite.
2. test_system_coverage_gate.py — Critical Path Coverage Enforcement:
Reads coverage_e2e.json and enforces minimum coverage thresholds on
production-critical modules. Each threshold is tied to a real production
incident. If persistent_list.py drops below 20% coverage, the build
fails — because that's exactly how the sessions.json corruption
bug went undetected.
Usage: pytest tests/e2e --cov=GramAddict.core \
--cov-report=json:coverage_e2e.json
Critical paths enforced:
- session_state.py (>=50%): serialization crash prevention
- persistent_list.py (>=20%): persist() path must be tested
- dopamine_engine.py (>=40%): session timeout logic
- screen_identity.py (>=60%): screen identification
- spatial_parser.py (>=70%): UI element parsing
- intent_resolver.py (>=50%): click decision logic
- q_nav_graph.py (>=25%): navigation integrity
- device_facade.py (>=40%): device interface
Replaces reactive bug-specific tests with 4 structural CONTRACTS
that make entire categories of bugs impossible:
1. Serialization Contract (13 parametrized hostile payloads):
Proves SessionStateEncoder NEVER crashes regardless of what
gets injected onto args (datetime, lambda, bytes, inf, nan, etc.)
2. Persistence Contract:
Forces the REAL persist() path by removing PYTEST_CURRENT_TEST guard.
Validates write → read round-trip with poisoned sessions.
3. Production Parity Contract:
Scans ALL GramAddict source for pytest/test-mode divergence guards.
Any unaudited divergence point fails the build. New guards require
explicit justification in KNOWN_DIVERGENCES registry.
4. Import Integrity Contract:
Imports every single GramAddict module to catch syntax errors,
circular imports, and missing dependencies at test time.
This addresses the root systemic failure: tests were running in a
parallel universe where serialization was silently skipped, allowing
a datetime injection to corrupt sessions.json and kill production.
Root cause: bot_flow.py injected configs.args.global_start_time = datetime.now()
which polluted the args namespace. SessionStateEncoder blindly serialized
args.__dict__ via json.dump, which crashed mid-write on the datetime object,
leaving sessions.json truncated/corrupt. Every subsequent restart failed.
Fixes:
- Remove global_start_time from configs.args (only lives on engine class vars)
- Harden SessionStateEncoder with _sanitize_value() to convert any
non-JSON-serializable type (datetime, timedelta, arbitrary objects) to strings
- Add test_system_session_persistence.py with 4 tests covering the exact
production crash scenario (datetime injection → json.dumps → round-trip)
- Fix test_engine_timeout.py broken GoalExecutor module reference
Bug 1 — VLM Profile Tab Hallucination:
- VLM confused 'Profile' nav tab with post author username
- Added Author Tab Guard to filter_navigation_conflicts()
- Fixed mock LLM to use structural row_feed_photo_profile matching
instead of hardcoded username (was the test lie)
Bug 2 — Like Button 1-Byte Delta Kill:
- Toggle actions (like/save) produce 1-byte XML deltas
- GOAP's MIN_UI_CHANGE_BYTES=50 threshold killed them as 'no change'
- Split interaction gate: interactions use ANY xml diff, navigations
keep the 50-byte threshold
Bug 3 — Empty Username Silently Accepted:
- PostDataExtraction returned 'Post by @' with no warning
- Added 'username_missing' reliability flag to result dict
- Downstream consumers can now detect degraded data quality
Cleanup:
- Removed all debug print() statements from production code
- Replaced with structured logger.debug() calls
PRODUCTION FIX 2026-04-30 12:45:
The entire E2E suite was systematically lying by testing a shell of the bot:
1. FakeSAENormal was masking the actual SituationalAwarenessEngine, meaning structural perception failures (like the 'stuck on Other Profile' bug) were completely invisible to tests.
2. The e2e_cognitive_stack was hardcoded to return None for 80% of engines (NavGraph, ZeroEngine, Telepathic, Darwin, ActiveInference), causing plugins to silently bypass their core logic during tests.
3. get_screenshot_b64 and screenshot() returned None, preventing the VLM stack from even running structural fallback code.
Fix:
- Systematically stripped FakeSAENormal from all 29 E2E workflow tests.
- Built e2e_cognitive_stack_factory to instantiate the REAL Cognitive Stack (GrowthBrain, QNavGraph, ZeroLatencyEngine, etc.) just like in bot_flow.py.
- Patched ONLY the external LLM network requests in conftest.py via mock_llm_network_calls so that structural execution stays 100% real but avoids non-deterministic LLM timeouts.
- Injected a valid 1x1 black image for screenshots.
All 29 E2E tests now execute the true production logic path and pass in 49 seconds.
PRODUCTION BUG 2026-04-30 11:50:
VLM returned action_bar_button_back (desc='Back') for 'tap profile tab'
intent on OTHER_PROFILE screen. This caused account_switcher to navigate
AWAY from the profile instead of opening the account selector bottom
sheet, halting the session.
Root cause: _visual_discovery pre-filter did not exclude navigation
buttons (Back, Close) when the intent was for a tab element.
Fix:
- Added filter_navigation_conflicts() method to IntentResolver
- Guards tab intents by excluding nodes with 'back' in resource_id
or content_desc == 'Back'
- Does NOT apply for 'tap back button' or 'press back' intents
TDD:
- RED: test_intent_resolver_back_guard.py — 2 tests that call
filter_navigation_conflicts on real OTHER_PROFILE XML
- GREEN: filter_navigation_conflicts method + wired into _visual_discovery
- E2E: test_workflow_account_switch.py, test_workflow_vlm_tab_confusion.py
31/31 tests passing.