Research Preview · Android OS · Autonomous Agent · Closed-Loop Verification

MobiMind: An autonomous AI agent which can control phones for you

MobiMind is an autonomous mobile agent that translates natural-language goals into grounded Android actions. By compressing noisy accessibility trees by 97.1%, pairing deterministic node interaction with multimodal visual fallback, and verifying state transitions at each step, MobiMind executes complex multi-app tasks reliably without fragile coordinate heuristics.

97.1%
Token Compression

~92,000 raw tree tokens compressed into ~2,600 semantic tokens, slashing LLM prompt latency and inference cost.

457 → 171
Node Pruning

Strips empty layout containers and structural wrappers while preserving deterministic device touch coordinates.

Dual-Path
Tree + Vision Fallback

Primary execution via Android Accessibility Service; Moondream visual grounding for undescribed icons.

Closed-Loop
Outcome Verification

Validates post-interaction screen state to confirm progress or trigger automatic recovery routines.

01 · System Overview

Meet MobiMind: Autonomous Phone Control

MobiMind translates natural language requests into a validated sequence of interactions with native Android applications and mobile web pages.

Task Benchmark: We assigned MobiMind the goal: “Compare auto fares on Rapido and Uber from my location to PES Modern College and tell me which one is cheaper.” Over a 10-step autonomous run, the agent queried both mobile apps, extracted live quotes in real time (Rapido ₹284 vs. Uber ₹254), and verified that Uber was ₹29.60 cheaper—stopping safely before booking.
Pipeline · Stage A

Task Formulation & Context Assembly

Deconstructs high-level human goals into context-grounded, sequential mobile actions.

  • Goal Ingestion: Ingests natural-language intent (e.g. “Compare auto fares across Uber and Rapido”).
  • Context Fusion: Combines foreground window hierarchies, previous turn history, and persistent app memory.
  • Incremental Execution: Re-evaluates state after every sub-action (opening drawers, querying search, selecting ride options).
Pipeline · Stage B

Observation & Grounded Device Control

Extracts native Android view trees and dispatches physical touch events deterministically.

  • Non-Root Extraction: Dumps view bounds, text, and roles directly from the Android Accessibility Service.
  • Noise Filtering: Strips redundant layout containers, reducing prompt token overhead by 97.1%.
  • Dual Grounding: Resolves semantic IDs to physical coordinates, with Moondream visual fallback for custom Canvas views.
End-to-end task loop: High-level goals become grounded interactions through structured observation, semantic parsing, LLM decision making, physical execution, and verification.
01.1 · Intended Users & Voice Interaction

Everyday phone tasks, through voice

MobiMind is intended to make Android tasks more accessible to older adults, blind users and people less familiar with technology.

Sarvam speech-to-text (STT) converts a spoken goal into text. Sarvam text-to-speech (TTS) provides spoken task updates, helping users follow the agent’s progress.

Older adults
Describe a task naturally, with less typing and menu navigation.
Blind users
Give spoken instructions and hear updates without relying only on visual feedback.
People less familiar with technology
Express what they need in everyday language, without learning each app’s interface.
Voice input and spoken feedback support access to everyday Android tasks.
  1. 01
    Speak a goalSarvam STT · Speech → text
  2. 02
    Execute the taskMobiMind · Android actions
  3. 03
    Hear progress updatesSarvam TTS · Text → speech
02 · Recorded Demos

Recorded task demonstrations

Five Android runs with step-by-step replay, audio and execution logs. Select a screenshot to inspect it in full.

01 · Multi-app10 steps

Fare comparison

Uber identified as the cheaper option.

Replay demo
02 · Web navigation9 steps

Department HOD

AI & Data Science department head identified.

Replay demo
03 · Code execution7 steps

Nepal flag in C

C program compiled and flag rendered.

Replay demo
04 · Visual grounding11 steps

Moon photo editing

Moon photo located, adjusted and saved.

Replay demo
05 · Media control10 steps

YouTube playback

Requested song found and playback verified.

Replay demo
03 · Architecture

The Closed-Loop Task Engine

The system continuously cycles through observation, semantic processing, decision making, hardware execution, and verification until the task completes or requires user input.

Phase 01 · Task Execution Loop

01 · Goal & Trigger

User enters a natural-language goal through the non-intrusive floating assistive overlay bubble.

The user taps MobiMind’s floating bubble on Android and states an outcome (e.g. “Compare auto fares on Rapido and Uber”). The mobile client creates a session token, captures foreground app identity, and connects to the FastAPI reasoning server.

Input User Voice / Text Prompt + Foreground Package
Output Initialized Session (Task ID: silver_wren)
Reliability & Safety Interactive confirmation dialogs for high-risk irreversible actions
// Phase 01: Goal Ingestion & Trigger
{
  "task_id": "silver_wren",
  "goal": "Compare auto fares on Rapido and Uber to PES Modern College",
  "foreground_app": "com.rapido.passenger"
}
Edge Client · Mobile Execution

Android Device Client

Tech Stack: Kotlin · Android SDK 34 · AccessibilityService · MediaProjection API · OkHttp / Retrofit · Coroutines · System Alert Window

  • Non-Root Accessibility Capture: Uses Android’s AccessibilityService to dump live window nodes, text labels, bounds, and clickable states directly from OS memory with zero root requirements.
  • Lossless Visual Capture: MediaProjection API captures live frame buffers on demand for Moondream visual grounding when controls lack text descriptions.
  • Floating Assistive Bubble: System Alert Window renders a persistent, non-intrusive assistive bubble for user goal input, execution progress feedback, and manual override confirmation.
  • Hardware Action Dispatcher: Dispatches physical interactions via accessibility node actions (e.g. ACTION_CLICK, ACTION_SET_TEXT) or ADB gesture taps (dispatchGesture), reporting immediate OS-level execution status.
  • Production Use Case: Serves as the eyes and hands on the user’s mobile device, transmitting sanitized screen state to the server and executing physical gestures safely.
Brain Engine · Cloud / Server

Reasoning Server Engine

Tech Stack: Python 3.11+ · FastAPI (ASGI / Uvicorn) · Pydantic v2 · LiteLLM / Gemini 2.5 / Claude 3.5 / GPT-4o · Moondream Vision · PyYAML

  • Semantic Tree Compression: Strips redundant layout wrappers, empty FrameLayouts, and non-interactive metadata, slashing raw tree size by 97.1% (from ~92,000 to ~2,600 tokens).
  • Dynamic Context & Memory Assembly: Fuses the user’s goal, recent turn history, active window hierarchy, and 3-tier persistent memory into model prompts.
  • Dual-Path Action Grounding: Resolves model decisions back to hardware coordinates; calls Moondream visual grounding if target icons lack accessible text.
  • Closed-Loop Outcome Verification: Computes state diffs between consecutive observations to verify task progress, detecting stalls or unexpected dialogs and triggering automated self-healing.
  • Production Use Case: Centralizes LLM decision making, API key security, prompt caching, telemetry serialization, and interactive web dashboard session replays.
Closed-loop architecture: Every selected action is connected to the subsequent screen capture. Verification creates context for continuing or self-healing.
04 · Screen Representation

The Android Accessibility Tree

An accessibility tree is a hierarchy of interface elements exposed by an application to accessibility services. Each node describes a UI component and its spatial relationship to other elements.

A node can represent a button, text field, image, list item, or layout container. Android exposes properties such as text, content description, class name, bounds, state, and supported actions through AccessibilityNodeInfo.

Parent–child relationships describe containment. For example, a product card contains a product name, price tag, and Add button. Capturing these relationships associates an action with the item it belongs to.

Available Screen Data & Edge Cases

The tree contains only what the application exposes. An icon may have an informative description, a blank description, or no node if rendered in custom Canvas. MobiMind captures raw nodes with parent IDs and window metadata to distinguish apps from permission dialog overlays.

Common Raw Accessibility Properties Used in Automation
Property Meaning Use in MobiMind Automation
node_id / parent_id Node identity and containment hierarchy. Associates child content and resolves physical device touch targets.
text / content_description Visible text or accessible label. Identifies Menu, Search, Add, or named UI buttons.
class_name / role Android element class type. Distinguishes inputs, buttons, links, and layout containers.
bounds Screen coordinates [x1, y1, x2, y2]. Grounds physical touch coordinates when gestures are executed.
enabled / checked / selected Current control state. Prevents clicking disabled controls and checks toggle changes.
actions Operations advertised by the node. Determines whether click, text entry, or scrolling is supported.
Accessibility mapping: Search fields, suggestion rows, content cards, and navigation controls mapped to a structured Android hierarchy.
Core Concept · UI Hierarchy

What is an Accessibility Tree? Android’s Native DOM

If you are familiar with web development, you already understand how mobile accessibility trees work. Just as a web browser parses HTML into a Document Object Model (DOM) of nested <div>, <button>, and <a> elements with styles and event listeners, the Android operating system maintains a parallel Accessibility Tree of AccessibilityNodeInfo objects for native applications.

Website DOM

HTMLDocument Hierarchy
  • Nodes & Tags: Composed of semantic tags such as <button>, <input>, <h1>, and layout <div>s.
  • Attributes & Labels: Identified via id, class, aria-label, placeholder, and inner text.
  • Spatial Geometry: CSS box model computes element bounding rects (top, left, width, height).
  • Action Dispatch: JavaScript listeners handle interaction (click(), dispatchEvent()).

Android Accessibility Tree

AccessibilityNodeInfo Hierarchy
  • Nodes & Classes: Composed of view classes such as android.widget.Button, EditText, TextView, and FrameLayout.
  • Properties & Labels: Identified via text, contentDescription, viewIdResourceName, and hintText.
  • Hardware Bounds: Absolute screen display coordinates [left, top, right, bottom] reported directly by the Android window manager.
  • Action Dispatch: Native accessibility actions (ACTION_CLICK, ACTION_SET_TEXT, ACTION_SCROLL_FORWARD).

Why MobiMind Controls Phones via Accessibility Trees Instead of Raw Pixels:

Screenshot-only vision agents are slow, expensive, and fragile—they frequently miss small buttons, hallucinate touch coordinates, or misread text through OCR. By reading Android's native Accessibility Tree, MobiMind receives exact mathematical bounds, verified text labels, and live component states (clickable, enabled, focused) directly from the OS, bridging web browser automation reliability to mobile apps.

Structural Equivalence: Comparing web HTML DOM elements (left) with Android's native AccessibilityNodeInfo hierarchy (right). Both expose hierarchical UI trees with labels, bounds, and action handlers, allowing MobiMind to automate mobile apps deterministically.
05 · Deep Dive

Raw Accessibility Tree vs. Semantic Tree

Why send ~92,000 tokens of noisy XML when the model only needs ~2,600 tokens of semantic data? See exactly how MobiMind prunes raw trees into high-density tokens while preserving hardware touch targets.

Raw Accessibility Tree (Before)

Verbose Layout Bloat

Raw Android captures repeat package names (com.android.chrome), class names (android.widget.FrameLayout), nested layout wrappers, empty metadata fields, and coordinate arrays for every node.

  • Hundreds of non-interactive layout containers
  • Over 92,000 tokens per single screen turn
  • Higher input token cost (see the 10-step task estimates below)
  • Slow prompt serialization latency
MobiMind Semantic Tree (After)

Dense Semantic Components

The cleaner strips structural wrappers and assigns compact semantic IDs (e.g., button_1, input_1). The server retains the hardware mapping while feeding the model concise tokens.

  • Preserves visible text, roles, states, and action flags
  • Reduces tokens by 97.1% (down to ~2,600 tokens/step)
  • Sub-cent per step execution economics
  • Deterministic server-side hardware grounding
Semantic compression: Raw accessibility trees converted into compact semantic components while retaining 1:1 execution coordinates on the reasoning server.
Interactive On-Page Comparison Playground

Compare Real Screen Data: Raw vs. Semantic

Recorded Chrome session: College Faculty Page · Step 005

Raw Nodes: 457
Semantic Nodes: 171
Compression: 97.1%
Raw Accessibility Tree (~92,000 tokens) 457 Nodes
Loading raw session capture…
Semantic Model Input (~2,600 tokens) 171 Elements
Loading semantic representation…
05.1 · Computational Cost

Task cost and capacity

Estimated inference cost for one 10-step task, before and after semantic compression.

Table · Cost per task (USD) and whole tasks per $1
Model Input priceUSD / 1M tokens Output priceUSD / 1M tokens Raw treeCost / task Semantic treeCost / task Tasks / $1Semantic tree
GPT-6 Luna $0.10 $0.50 $0.09240 $0.00300 333
Claude Haiku 5.5 $0.10 $0.50 $0.09240 $0.00300 333
Gemini 3.8 FlashPromotional rate $0.75 $3.75 $0.69300 $0.02250 44
DeepSeek V4.1 FlashPeak rate $0.30 $1.20 $0.27696 $0.00876 114
DeepSeek V4.1 FlashOff-peak rate $0.15 $0.60 $0.13848 $0.00438 228
Ctask = TinPin + ToutPout1,000,000 N$1 = ⌊$1Ctask⌋

T: total tokens per task · P: price per million tokens · ⌊ ⌋: round down

Example · GPT-6 Luna
$0.00260Input cost $0.00040Output cost $0.00300One task

$1 ÷ $0.00300 → 333 whole tasks

Assumptions: 10 calls per task; 92,000 raw or 2,600 semantic input tokens and 80 billed output tokens per call. Standard uncached text rates, checked 8 October 2026. Additional context, reasoning, retries and services are excluded.

Rates apply below long-context thresholds. Gemini’s promotional rate ends 31 December 2026. DeepSeek varies by peak/off-peak hours. Figures estimate budget capacity, not task success.

06 · Visual Grounding

Finding Missing Targets: Moondream Fallback

The agent must determine whether a target is offscreen, hidden in a drawer, omitted during compression, or only identifiable from the screenshot.

Missing-Target Cases and Automated Recovery Steps
Case Real-World Example Automated Next Step
Target is offscreen Product appears below visible viewport. Scrolls supported region, re-captures tree, and uses new target ID.
Target is hidden Link is inside a collapsed hamburger drawer. Taps Menu button and inspects newly exposed navigation links.
Control has no node Map recenter icon drawn directly on custom Canvas. Sends target description + screenshot to Moondream visual grounding.
Node lacks label Photo gallery with empty ImageView descriptions. Uses visual grounding to identify the requested image subject.
Ambiguous targets Multiple products with identical names. Pauses execution to query the user for clarifying details.
Visual grounding fallback: Locating an undescribed map recenter control and distinguishing specific photos using multimodal visual grounding.
Example Visual Target Request Schema
{
  "type": "visual_tap",
  "target": "the map recenter icon at the lower right",
  "tool": "moondream_grounding",
  "action": "tap"
}
07 · Decision & Execution

From Decisions to Physical Android Touches

The model selects an action and a current semantic target. The server validates that choice and resolves it to executable device coordinates.

{
  "type": "tap",
  "target": {
    "id": "button_1",
    "label": "Toggle navigation"
  },
  "rationale": "Menu drawer must be opened to locate faculty link"
}
Decision interface: User goal, active memory, previous turns, semantic components, and optional visual evidence inform exactly one next structured decision.
Hardware execution: Grounding connects semantic decisions to physical interactions via accessibility click, touch coordinates, and escalation.
08 · Reliability & Recovery

Closed-Loop Verification and Recovery

Execution feedback reports the command dispatch result. Verification evaluates whether the new observation satisfies the expected state transition.

Failure Mode 01

Unchanged Screen

An accepted touch leaves the screen unchanged due to misaligned coordinates. MobiMind re-grounds the target with direct coordinates or retries with a slight offset.

Failure Mode 02

Unexpected Destination

The opened page does not match expected post-conditions. The agent inspects navigation results and triggers corrective navigation instead of falsely marking completion.

Failure Mode 03

Unsaved / Uncommitted State

A field displays edited text, but persistence is required. The agent saves and verifies the stored value before reporting success.

Recovery flow: Recovering from interaction failures through re-observation, re-matching, grounded touch escalation, or alternative route selection.
09 · Knowledge Persistence

Contextual Memory Architecture

MobiMind separates reusable cross-task knowledge from the temporary execution state of an active task.

Tier 01 · Global Scope

User Preferences

Information that applies across applications, such as preferred language and default addresses.

Tier 02 · Package Scope

Application Mappings

Friendly application names associated with Android package identifiers (e.g. “Rapido” → com.rapido.passenger).

Tier 03 · App Scope

Application Knowledge

Navigation landmarks, known animation delays, and reusable observations for specific apps.

Memory hierarchy: Separating persistent knowledge from ephemeral checklist state.
10 · Security & User Control

User control and authentication

Ask before acting

MobiMind pauses for approval when required and asks for clarification when an instruction is unclear.

Confirmation · Permission to act

“Book an Uber Auto from your current location to PES Modern College for ₹254?”

Confirm booking / Cancel
The ride is booked only after confirmation.

Ask user · Missing information

“Which Priya: Priya Shah or Priya Rao?”

User: “Priya Shah.”
Selecting a contact does not approve sending.

Illustrative examples. Approval expires and applies only to the action shown.

RSA authentication

Each party signs messages with its private key. The receiver checks the signature using the sender’s trusted public key, verifying the sender and detecting changes.

App signs → server verifies → server signs → app verifies.

RSA signing is proposed; approval checks are implemented. Signatures verify messages; encryption protects their contents.

11 · Research & Evaluation

Research Context & Mobile Agent Benchmarks

How MobiMind compares against prominent mobile agent frameworks in the research literature: AutoDroid, AndroidWorld, and M3A.

Architectural positioning: Comparing screenshot-only agents, raw accessibility parsers, and MobiMind’s hybrid semantic representation.
Academic Reference

AutoDroid (Wen et al., 2023)

Research on Android task automation using language models, UI representations, and app-specific memory. MobiMind extends this by adding continuous closed-loop post-action verification and multimodal visual fallback.

Read AutoDroid on arXiv ↗
Academic Reference

AndroidWorld (Rawles et al., 2024)

A mobile-agent benchmark featuring parameterized tasks, reproducible initialization, and reward verification. Guides MobiMind’s outcome validation philosophy.

Read AndroidWorld on arXiv ↗
Interactive Experience

Inspect Real Execution Traces

Explore complete session recordings or examine the exact JSON screen structures processed by the model.