MobiMind: An autonomous AI agent which can control phones for you
MobiMind is an autonomous mobile agent that translates natural-language goals into grounded Android actions. By compressing noisy accessibility trees by 97.1%, pairing deterministic node interaction with multimodal visual fallback, and verifying state transitions at each step, MobiMind executes complex multi-app tasks reliably without fragile coordinate heuristics.
~92,000 raw tree tokens compressed into ~2,600 semantic tokens, slashing LLM prompt latency and inference cost.
Strips empty layout containers and structural wrappers while preserving deterministic device touch coordinates.
Primary execution via Android Accessibility Service; Moondream visual grounding for undescribed icons.
Validates post-interaction screen state to confirm progress or trigger automatic recovery routines.
Meet MobiMind: Autonomous Phone Control
MobiMind translates natural language requests into a validated sequence of interactions with native Android applications and mobile web pages.
Task Formulation & Context Assembly
Deconstructs high-level human goals into context-grounded, sequential mobile actions.
- Goal Ingestion: Ingests natural-language intent (e.g. “Compare auto fares across Uber and Rapido”).
- Context Fusion: Combines foreground window hierarchies, previous turn history, and persistent app memory.
- Incremental Execution: Re-evaluates state after every sub-action (opening drawers, querying search, selecting ride options).
Observation & Grounded Device Control
Extracts native Android view trees and dispatches physical touch events deterministically.
- Non-Root Extraction: Dumps view bounds, text, and roles directly from the Android Accessibility Service.
- Noise Filtering: Strips redundant layout containers, reducing prompt token overhead by 97.1%.
- Dual Grounding: Resolves semantic IDs to physical coordinates, with Moondream visual fallback for custom Canvas views.
Everyday phone tasks, through voice
MobiMind is intended to make Android tasks more accessible to older adults, blind users and people less familiar with technology.
Sarvam speech-to-text (STT) converts a spoken goal into text. Sarvam text-to-speech (TTS) provides spoken task updates, helping users follow the agent’s progress.
- Older adults
- Describe a task naturally, with less typing and menu navigation.
- Blind users
- Give spoken instructions and hear updates without relying only on visual feedback.
- People less familiar with technology
- Express what they need in everyday language, without learning each app’s interface.
- 01Speak a goalSarvam STT · Speech → text
- 02Execute the taskMobiMind · Android actions
- 03Hear progress updatesSarvam TTS · Text → speech
Recorded task demonstrations
Five Android runs with step-by-step replay, audio and execution logs. Select a screenshot to inspect it in full.
The Closed-Loop Task Engine
The system continuously cycles through observation, semantic processing, decision making, hardware execution, and verification until the task completes or requires user input.
01 · Goal & Trigger
User enters a natural-language goal through the non-intrusive floating assistive overlay bubble.
The user taps MobiMind’s floating bubble on Android and states an outcome (e.g. “Compare auto fares on Rapido and Uber”). The mobile client creates a session token, captures foreground app identity, and connects to the FastAPI reasoning server.
// Phase 01: Goal Ingestion & Trigger
{
"task_id": "silver_wren",
"goal": "Compare auto fares on Rapido and Uber to PES Modern College",
"foreground_app": "com.rapido.passenger"
}
Android Device Client
Tech Stack: Kotlin · Android SDK 34 · AccessibilityService · MediaProjection API · OkHttp / Retrofit · Coroutines · System Alert Window
- Non-Root Accessibility Capture: Uses Android’s AccessibilityService to dump live window nodes, text labels, bounds, and clickable states directly from OS memory with zero root requirements.
- Lossless Visual Capture: MediaProjection API captures live frame buffers on demand for Moondream visual grounding when controls lack text descriptions.
- Floating Assistive Bubble: System Alert Window renders a persistent, non-intrusive assistive bubble for user goal input, execution progress feedback, and manual override confirmation.
-
Hardware Action Dispatcher: Dispatches physical interactions via accessibility node actions (e.g.
ACTION_CLICK,ACTION_SET_TEXT) or ADB gesture taps (dispatchGesture), reporting immediate OS-level execution status. - Production Use Case: Serves as the eyes and hands on the user’s mobile device, transmitting sanitized screen state to the server and executing physical gestures safely.
Reasoning Server Engine
Tech Stack: Python 3.11+ · FastAPI (ASGI / Uvicorn) · Pydantic v2 · LiteLLM / Gemini 2.5 / Claude 3.5 / GPT-4o · Moondream Vision · PyYAML
- Semantic Tree Compression: Strips redundant layout wrappers, empty FrameLayouts, and non-interactive metadata, slashing raw tree size by 97.1% (from ~92,000 to ~2,600 tokens).
- Dynamic Context & Memory Assembly: Fuses the user’s goal, recent turn history, active window hierarchy, and 3-tier persistent memory into model prompts.
- Dual-Path Action Grounding: Resolves model decisions back to hardware coordinates; calls Moondream visual grounding if target icons lack accessible text.
- Closed-Loop Outcome Verification: Computes state diffs between consecutive observations to verify task progress, detecting stalls or unexpected dialogs and triggering automated self-healing.
- Production Use Case: Centralizes LLM decision making, API key security, prompt caching, telemetry serialization, and interactive web dashboard session replays.
The Android Accessibility Tree
An accessibility tree is a hierarchy of interface elements exposed by an application to accessibility services. Each node describes a UI component and its spatial relationship to other elements.
A node can represent a button, text field, image, list item, or layout container. Android exposes properties such as text, content description, class name, bounds, state, and supported actions through AccessibilityNodeInfo.
Parent–child relationships describe containment. For example, a product card contains a product name, price tag, and Add button. Capturing these relationships associates an action with the item it belongs to.
Available Screen Data & Edge Cases
The tree contains only what the application exposes. An icon may have an informative description, a blank description, or no node if rendered in custom Canvas. MobiMind captures raw nodes with parent IDs and window metadata to distinguish apps from permission dialog overlays.
| Property | Meaning | Use in MobiMind Automation |
|---|---|---|
node_id / parent_id |
Node identity and containment hierarchy. | Associates child content and resolves physical device touch targets. |
text / content_description |
Visible text or accessible label. | Identifies Menu, Search, Add, or named UI buttons. |
class_name / role |
Android element class type. | Distinguishes inputs, buttons, links, and layout containers. |
bounds |
Screen coordinates [x1, y1, x2, y2]. | Grounds physical touch coordinates when gestures are executed. |
enabled / checked / selected |
Current control state. | Prevents clicking disabled controls and checks toggle changes. |
actions |
Operations advertised by the node. | Determines whether click, text entry, or scrolling is supported. |
What is an Accessibility Tree? Android’s Native DOM
If you are familiar with web development, you already understand how mobile accessibility trees work. Just as a web browser parses HTML into a Document Object Model (DOM) of nested <div>, <button>, and <a> elements with styles and event listeners, the Android operating system maintains a parallel Accessibility Tree of AccessibilityNodeInfo objects for native applications.
Website DOM
HTMLDocument Hierarchy-
Nodes & Tags: Composed of semantic tags such as
<button>,<input>,<h1>, and layout<div>s. -
Attributes & Labels: Identified via
id,class,aria-label,placeholder, and inner text. -
Spatial Geometry: CSS box model computes element bounding rects (
top,left,width,height). -
Action Dispatch: JavaScript listeners handle interaction (
click(),dispatchEvent()).
Android Accessibility Tree
AccessibilityNodeInfo Hierarchy-
Nodes & Classes: Composed of view classes such as
android.widget.Button,EditText,TextView, andFrameLayout. -
Properties & Labels: Identified via
text,contentDescription,viewIdResourceName, andhintText. -
Hardware Bounds: Absolute screen display coordinates
[left, top, right, bottom]reported directly by the Android window manager. -
Action Dispatch: Native accessibility actions (
ACTION_CLICK,ACTION_SET_TEXT,ACTION_SCROLL_FORWARD).
Why MobiMind Controls Phones via Accessibility Trees Instead of Raw Pixels:
Screenshot-only vision agents are slow, expensive, and fragile—they frequently miss small buttons, hallucinate touch coordinates, or misread text through OCR. By reading Android's native Accessibility Tree, MobiMind receives exact mathematical bounds, verified text labels, and live component states (clickable, enabled, focused) directly from the OS, bridging web browser automation reliability to mobile apps.
Raw Accessibility Tree vs. Semantic Tree
Why send ~92,000 tokens of noisy XML when the model only needs ~2,600 tokens of semantic data? See exactly how MobiMind prunes raw trees into high-density tokens while preserving hardware touch targets.
Verbose Layout Bloat
Raw Android captures repeat package names (com.android.chrome), class names (android.widget.FrameLayout), nested layout wrappers, empty metadata fields, and coordinate arrays for every node.
- Hundreds of non-interactive layout containers
- Over 92,000 tokens per single screen turn
- Higher input token cost (see the 10-step task estimates below)
- Slow prompt serialization latency
Dense Semantic Components
The cleaner strips structural wrappers and assigns compact semantic IDs (e.g., button_1, input_1). The server retains the hardware mapping while feeding the model concise tokens.
- Preserves visible text, roles, states, and action flags
- Reduces tokens by 97.1% (down to ~2,600 tokens/step)
- Sub-cent per step execution economics
- Deterministic server-side hardware grounding
Compare Real Screen Data: Raw vs. Semantic
Recorded Chrome session: College Faculty Page · Step 005
Loading raw session capture…
Loading semantic representation…
Task cost and capacity
Estimated inference cost for one 10-step task, before and after semantic compression.
| Model | Input priceUSD / 1M tokens | Output priceUSD / 1M tokens | Raw treeCost / task | Semantic treeCost / task | Tasks / $1Semantic tree |
|---|---|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.50 | $0.09240 | $0.00300 | 333 |
| Claude Haiku 5.5 | $0.10 | $0.50 | $0.09240 | $0.00300 | 333 |
| Gemini 3.8 FlashPromotional rate | $0.75 | $3.75 | $0.69300 | $0.02250 | 44 |
| DeepSeek V4.1 FlashPeak rate | $0.30 | $1.20 | $0.27696 | $0.00876 | 114 |
| DeepSeek V4.1 FlashOff-peak rate | $0.15 | $0.60 | $0.13848 | $0.00438 | 228 |
T: total tokens per task · P: price per million tokens · ⌊ ⌋: round down
$1 ÷ $0.00300 → 333 whole tasks
Assumptions: 10 calls per task; 92,000 raw or 2,600 semantic input tokens and 80 billed output tokens per call. Standard uncached text rates, checked 8 October 2026. Additional context, reasoning, retries and services are excluded.
Rates apply below long-context thresholds. Gemini’s promotional rate ends 31 December 2026. DeepSeek varies by peak/off-peak hours. Figures estimate budget capacity, not task success.
Finding Missing Targets: Moondream Fallback
The agent must determine whether a target is offscreen, hidden in a drawer, omitted during compression, or only identifiable from the screenshot.
| Case | Real-World Example | Automated Next Step |
|---|---|---|
| Target is offscreen | Product appears below visible viewport. | Scrolls supported region, re-captures tree, and uses new target ID. |
| Target is hidden | Link is inside a collapsed hamburger drawer. | Taps Menu button and inspects newly exposed navigation links. |
| Control has no node | Map recenter icon drawn directly on custom Canvas. | Sends target description + screenshot to Moondream visual grounding. |
| Node lacks label | Photo gallery with empty ImageView descriptions. | Uses visual grounding to identify the requested image subject. |
| Ambiguous targets | Multiple products with identical names. | Pauses execution to query the user for clarifying details. |
{
"type": "visual_tap",
"target": "the map recenter icon at the lower right",
"tool": "moondream_grounding",
"action": "tap"
}
From Decisions to Physical Android Touches
The model selects an action and a current semantic target. The server validates that choice and resolves it to executable device coordinates.
{
"type": "tap",
"target": {
"id": "button_1",
"label": "Toggle navigation"
},
"rationale": "Menu drawer must be opened to locate faculty link"
}
Closed-Loop Verification and Recovery
Execution feedback reports the command dispatch result. Verification evaluates whether the new observation satisfies the expected state transition.
Unchanged Screen
An accepted touch leaves the screen unchanged due to misaligned coordinates. MobiMind re-grounds the target with direct coordinates or retries with a slight offset.
Unexpected Destination
The opened page does not match expected post-conditions. The agent inspects navigation results and triggers corrective navigation instead of falsely marking completion.
Unsaved / Uncommitted State
A field displays edited text, but persistence is required. The agent saves and verifies the stored value before reporting success.
Contextual Memory Architecture
MobiMind separates reusable cross-task knowledge from the temporary execution state of an active task.
User Preferences
Information that applies across applications, such as preferred language and default addresses.
Application Mappings
Friendly application names associated with Android package identifiers (e.g. “Rapido” → com.rapido.passenger).
Application Knowledge
Navigation landmarks, known animation delays, and reusable observations for specific apps.
User control and authentication
Ask before acting
MobiMind pauses for approval when required and asks for clarification when an instruction is unclear.
Confirmation · Permission to act
“Book an Uber Auto from your current location to PES Modern College for ₹254?”
Confirm booking / Cancel
The ride is booked only after confirmation.
Ask user · Missing information
“Which Priya: Priya Shah or Priya Rao?”
User: “Priya Shah.”
Selecting a contact does not approve sending.
Illustrative examples. Approval expires and applies only to the action shown.
RSA authentication
Each party signs messages with its private key. The receiver checks the signature using the sender’s trusted public key, verifying the sender and detecting changes.
RSA signing is proposed; approval checks are implemented. Signatures verify messages; encryption protects their contents.
Research Context & Mobile Agent Benchmarks
How MobiMind compares against prominent mobile agent frameworks in the research literature: AutoDroid, AndroidWorld, and M3A.
AutoDroid (Wen et al., 2023)
Research on Android task automation using language models, UI representations, and app-specific memory. MobiMind extends this by adding continuous closed-loop post-action verification and multimodal visual fallback.
Read AutoDroid on arXiv ↗AndroidWorld (Rawles et al., 2024)
A mobile-agent benchmark featuring parameterized tasks, reproducible initialization, and reward verification. Guides MobiMind’s outcome validation philosophy.
Read AndroidWorld on arXiv ↗Inspect Real Execution Traces
Explore complete session recordings or examine the exact JSON screen structures processed by the model.