Uber or Rapido?
Follow the agent as it checks both Auto fares and compares the results.
View fare comparisonMobiMind uses Android accessibility data to identify interface elements, selects one action with a language model, and checks the resulting screen.
This page explains the accessibility tree, semantic node processing, action execution, and visual fallback. Interactive examples show the data before and after processing.
Replay real tasks with screenshots, recorded voice, and the agent’s decisions.
Follow the agent as it checks both Auto fares and compares the results.
View fare comparisonTrace the route from a college homepage to the AI & Data Science faculty page.
View website researchWatch the recorded steps for entering a C program and checking its rendered output.
View coding demoRead-only recordings · five selected runs · no device connection
MobiMind is an Android automation prototype that converts a user request into a sequence of interactions with applications and browser pages.
The user specifies an outcome, such as opening a web page, selecting a product, or changing a setting. The agent uses the foreground application, current screen, task history, and relevant memory to determine the next step.
The task is executed incrementally. A link may require opening a drawer; a product may require searching and selecting a variant. These intermediate states become new inputs to the decision process.
The Android client captures interface data and executes device commands. The server processes accessibility nodes, builds the model input, validates the decision, and resolves its target.
When structural data cannot identify a visible target, a screenshot can support visual grounding. The agent then checks whether the interaction produced the expected application state.
The system repeats observation, processing, decision, execution, and verification until the goal is completed, user input is required, or execution is blocked.
Capture the app, accessibility nodes, windows, and optional screenshot.
Create semantic components and retain execution mappings.
Request one structured next action from the language model.
Validate the target and send the command to Android.
Compare the new observation with the expected result.
The client owns accessibility capture, screenshot capture, the task interface, local memory, and device actions. It reports whether the requested command was accepted or failed.
The server owns screen interpretation, model requests, output validation, target resolution, and session state. Command acceptance is recorded separately from evidence that the task progressed.
An accessibility tree is a hierarchy of interface elements exposed by an application to accessibility services. Each node describes an element and its relationship to other elements.
A node can represent a button, text field, image, list item, or layout container. Android exposes properties such as text, content description, class, bounds, state, and supported actions through AccessibilityNodeInfo.
Parent–child relationships describe containment. For example, a product card may contain a product name, pack size, price, and Add button. Capturing these relationships helps associate an action with the item it belongs to.
The tree contains information provided by the application. An icon may have a useful description, an empty description, or no separate node. A screenshot may therefore contain visible information that is absent from the tree.
MobiMind records nodes with parent IDs and window metadata. Window information helps distinguish an application from a permission dialog or another overlay above it.
| Property | Meaning | Use in automation |
|---|---|---|
node_id / parent_id | Node identity and containment. | Associate child content and resolve device targets. |
text / content_description | Visible text or an accessible label. | Identify Menu, Search, Add, or a named item. |
class_name / role | Element type. | Distinguish inputs, buttons, links, and containers. |
bounds | Position and dimensions on the screen. | Ground a touch when geometry is needed. |
enabled / checked / selected | Current control state. | Avoid disabled actions and check selection changes. |
actions | Operations advertised by the node. | Determine whether click, text entry, or scrolling is supported. |
Native applications expose interface elements through Android’s accessibility framework. Browsers can expose web roles and links through the same framework. MobiMind processes these two surfaces with separate cleaners and converts their output into a shared semantic representation. The current browser pipeline uses Android accessibility captures; it does not retrieve the page DOM directly.
Semantic nodes are processed components derived from raw accessibility data. They describe meaningful content and controls using compact labels, roles, state, and actions.
A raw capture repeats package names, class names, flags, empty fields, and layout information across many nodes. Sending every property to the model increases the input size without necessarily improving target selection.
The cleaner identifies useful elements, consolidates content where appropriate, and removes redundant structure from the model-facing representation. Each retained component receives a semantic ID, such as input_1 or button_1.
The server keeps component data, raw node mappings, action mappings, and geometry. The model reads compact text containing semantic IDs and supported actions. It does not need the full execution metadata to select an action.
A semantic ID identifies a target in a particular observation. After navigation or a layout change, the agent must use the IDs from the new capture.
Labels, roles, values, useful state, supported actions, and availability indicators help the model identify and operate controls.
Repeated metadata, empty fields, structural wrappers, and duplicate content can be omitted or consolidated in the compact model input.
Compression can omit useful information. The recorded example below shows the raw capture and the semantic text used in the same session step.
A Chrome capture of the college faculty page, taken from the session logs. Both representations belong to the same step.
367,645 raw JSON characters → 10,520 semantic text characters. Raw JSON is measured without formatting whitespace.
Source: 2 October 2026 · bright_wren · step 005.
Fewer input tokens can reduce the cost of a model call when the provider charges per token. The recorded example shows the text reduction and a simple token estimate.
Actual tokenization depends on the model, text, and serialization. The comparison excludes the system prompt, task context, history, screenshots, output tokens, caching, and any visual-grounding calls.
Measure input size together with target retention and task outcomes. A small representation is useful only if it preserves the information required to select the correct action.
For example, a product choice can be ambiguous if its variant or price is missing. Missing information should trigger another observation, visual inspection, or a user question rather than an unsupported selection.
The agent must determine whether the target is offscreen, hidden, omitted during processing, or only identifiable from the screenshot.
| Case | Example | Next step |
|---|---|---|
| Target is offscreen | A product appears below the visible list. | Scroll a supported region, capture again, and use the current target ID. |
| Target is hidden | Alumni is inside a closed navigation drawer. | Open Menu and inspect the newly visible links. |
| Visible control has no useful node | A map recenter icon is drawn inside a custom view. | Describe the target and ground it in the current screenshot. |
| Node lacks visual meaning | Gallery images have empty descriptions. | Use the screenshot to identify the requested image. |
| Useful text was omitted | A product attribute is present on screen but missing from compact output. | Inspect available visual evidence or obtain more information before making a price-dependent choice. |
| Target is absent or ambiguous | No matching item, or several equally plausible items. | Search or navigate further, ask the user, or report the blocker. |
MobiMind can send a target description and the current screenshot to the Moondream grounding component. The returned location supports a tap or another configured gesture.
For “Open the car photo,” the screenshot can distinguish a car from mountains or portraits even when the tree provides no image description. A map recenter control can be located by its icon and position.
The server rejects a semantic action when its selected target is absent from the current observation. Inventing a semantic ID does not create an executable control.
Visual coordinates also depend on the current layout. After scrolling, opening a dialog, or rotating the device, the agent needs fresh grounding. A successful gesture still requires an outcome check.
{
"type": "visual_tap",
"target": "the map recenter icon at the lower right",
"tool": "moondream_grounding",
"action": "tap"
}The target is a description, not an invented node ID. Grounding supplies the location before the device executes the touch.
The model selects an action and a current semantic target. The server validates that choice and resolves it to executable device information.
The decision input includes the goal, current semantic screen, relevant memory, session state, and recent actions. An optional image provides visual context.
The output requests one next step. Actions include tapping, entering text, scrolling, changing progress, launching an app, waiting, visual interaction, and asking the user.
{
"type": "tap",
"target": {
"id": "button_1",
"label": "Menu"
}
}This example is valid only when the current observation exposes button_1 with the label Menu. The full decision includes additional session and result fields.
The server uses the semantic ID to retrieve the raw node and action mapping. Validation checks target existence, applicable state, and action support. A text-entry step requires an appropriate input; a disabled button requires resolving its prerequisite. Android executes the resolved command using the configured mechanism.
Each example specifies the request, the relevant observations, the action sequence, and the evidence needed to establish completion.
Request: “Open the Alumni page.” The initial observation contains Menu but no Alumni link. Tap Menu, capture the drawer, select Alumni, and capture the destination.
Verification: Check the destination heading and available page content. An accepted menu tap is insufficient if the drawer did not open. A loading screen requires waiting and another capture.
Request: “Add one 500 ml Amul Milk to the cart.” Identify the search input, enter the query, and inspect the matching product and pack size before selecting Add.
Verification: Check the matching cart item or quantity change. If important attributes are missing from semantic output, inspect the screen before selecting. Several matching variants may require clarification.
Request: Delete a selected note. A confirmation dialog displays “Delete note?”, explanatory text, Cancel, and Delete. The agent must resolve the dialog controls from the new capture; background controls are not sufficient.
Verification: After the authorized confirmation, check that the dialog closed and the intended note is absent. Closing the dialog alone does not establish deletion. When user confirmation is required, hold the action until the user responds.
Request: “Open the car photo.” A gallery may expose ImageView nodes with empty descriptions. Those labels do not identify which image contains a car.
Verification: Ground the car thumbnail using the screenshot, tap it, and inspect the opened image. If multiple car photos are visible, ask for a distinguishing detail.
Request: “Set the slider to 55.” Read its exposed range and current value. Use a supported progress action, or ground the handle and track if the control requires dragging.
Verification: Read the resulting value. If the setting requires Save or Apply, perform that step only after verifying the adjustment, then check that the value persisted.
Execution feedback reports the command result. Verification evaluates whether the new observation satisfies the expected state change.
An accepted touch can leave the interface unchanged. Loading, an overlay, a stale target, or incorrect geometry can prevent the intended result.
The agent records the new state and classifies the result as verified, failed, or uncertain. Uncertainty means the evidence is insufficient; it should not be treated as completion.
Depending on the observation, recovery can include waiting, locating the target again, strengthening grounding, using a screenshot, or choosing a different navigation route.
Repeated attempts use bounded budgets. A retry should use current information and be followed by another outcome check. If execution cannot continue, the agent reports the blocker or requests user input.
After tapping Menu, the next capture still shows the closed page. Re-observe the control and retry with a valid target or stronger grounding.
The opened page does not match Alumni. Inspect the navigation result and choose an appropriate correction instead of marking the task complete.
A field displays an edited value, but persistence is required. Save the change and check the stored result before reporting completion.
MobiMind separates reusable knowledge from the temporary state of an active task.
The Android client stores memory in a Markdown file. Relevant sections accompany observations, and the server can return updates for persistence.
A saved navigation note may indicate that Alumni is inside the menu. It helps select a route, but the agent still needs to locate the current control and verify the destination.
Information that applies across apps, such as preferred language.
Friendly application names associated with Android package identifiers.
Navigation details and reusable observations associated with the current package.
The current checklist, recent actions, verified facts, and blockers describe the active task. They support the next decision without automatically becoming long-term memory. Browser website notes can be stored in the browser’s app section; separate per-domain retrieval is a possible extension.
The agent can request missing information, hold an action for confirmation, or stop when the task cannot proceed.
A question resolves information the screen does not provide, such as which of several products the user intended. The answer continues the existing task with its context.
When confirmation is required, the server holds the proposed action. Approval leads to a fresh observation and decision. The user can also stop a running task from the Android interface.
The following diagram describes an RSA signing extension for client–server messages. The sender would sign a message with a private key, and the receiver would verify its signature with the corresponding public key.
This diagram is a proposed design. Request and response signing are not implemented in the inspected active communication path.
The main research question is whether a compact semantic representation can reduce model input while preserving the information needed for reliable mobile actions.
Compare raw-node input, semantic input, and semantic input with visual fallback on matched tasks. Measure input tokens, retained target information, target-selection errors, and completed outcomes.
Test difficult cases such as unlabeled icons, hidden controls, repeated products, dialogs, and custom sliders. Missing labels and omitted attributes should be recorded as representation failures when they affect the task.
Use reproducible initial states and independent final-state checks. Measure task success, false completion, action count, recovery attempts, latency, and provider-reported token usage.
Keep model settings and task conditions consistent when comparing components. The examples on this page demonstrate processing behavior; they do not establish a system-wide success rate or cost reduction.
Research on Android task automation using language models, UI representations, and app-specific knowledge.
A mobile-agent benchmark with parameterized tasks and reproducible initialization and outcome evaluation.
Replay a recorded run or inspect the screen data behind an agent decision.