Android agent · Research prototype

Semantic UI Understanding for Android

MobiMind uses Android accessibility data to identify interface elements, selects one action with a language model, and checks the resulting screen.

This page explains the accessibility tree, semantic node processing, action execution, and visual fallback. Interactive examples show the data before and after processing.

01
Structured observationLabels, roles, state, and available actions.
02
Semantic representationCompact model input with execution mappings.
03
Outcome verificationCheck the screen after each interaction.
Recorded demos

See MobiMind in action

Replay real tasks with screenshots, recorded voice, and the agent’s decisions.

Compare across apps

Uber or Rapido?

Follow the agent as it checks both Auto fares and compares the results.

View fare comparison
Research on the web

Find the department HOD

Trace the route from a college homepage to the AI & Data Science faculty page.

View website research
Write and run code

Draw Nepal’s flag

Watch the recorded steps for entering a C program and checking its rendered output.

View coding demo

Read-only recordings · five selected runs · no device connection

TopicsArchitectureAccessibility treeNode processingMissing targetsActionsRecoveryMemory
01 · Overview

Meet MobiMind

MobiMind is an Android automation prototype that converts a user request into a sequence of interactions with applications and browser pages.

Recorded task: MobiMind compared Auto quotes across Rapido and Uber and reported Uber as ₹29.60 cheaper. No ride was booked.

Task input

The user specifies an outcome, such as opening a web page, selecting a product, or changing a setting. The agent uses the foreground application, current screen, task history, and relevant memory to determine the next step.

The task is executed incrementally. A link may require opening a drawer; a product may require searching and selecting a variant. These intermediate states become new inputs to the decision process.

Observation and control

The Android client captures interface data and executes device commands. The server processes accessibility nodes, builds the model input, validates the decision, and resolves its target.

When structural data cannot identify a visible target, a screenshot can support visual grounding. The agent then checks whether the interaction produced the expected application state.

Natural-language requests become grounded interactions through understanding, decision making, execution, and checking the result. These are conceptual task examples.
02 · Architecture

The task loop

The system repeats observation, processing, decision, execution, and verification until the goal is completed, user input is required, or execution is blocked.

01Observe

Capture the app, accessibility nodes, windows, and optional screenshot.

02Process

Create semantic components and retain execution mappings.

03Decide

Request one structured next action from the language model.

04Execute

Validate the target and send the command to Android.

05Verify

Compare the new observation with the expected result.

Android client

The client owns accessibility capture, screenshot capture, the task interface, local memory, and device actions. It reports whether the requested command was accepted or failed.

Reasoning server

The server owns screen interpretation, model requests, output validation, target resolution, and session state. Command acceptance is recorded separately from evidence that the task progressed.

The architecture connects each selected action with the next observation. A menu-opening example shows how verification creates context for continuing.
03 · Screen representation

The accessibility tree

An accessibility tree is a hierarchy of interface elements exposed by an application to accessibility services. Each node describes an element and its relationship to other elements.

A node can represent a button, text field, image, list item, or layout container. Android exposes properties such as text, content description, class, bounds, state, and supported actions through AccessibilityNodeInfo.

Parent–child relationships describe containment. For example, a product card may contain a product name, pack size, price, and Add button. Capturing these relationships helps associate an action with the item it belongs to.

Available screen data

The tree contains information provided by the application. An icon may have a useful description, an empty description, or no separate node. A screenshot may therefore contain visible information that is absent from the tree.

MobiMind records nodes with parent IDs and window metadata. Window information helps distinguish an application from a permission dialog or another overlay above it.

Common raw accessibility properties
PropertyMeaningUse in automation
node_id / parent_idNode identity and containment.Associate child content and resolve device targets.
text / content_descriptionVisible text or an accessible label.Identify Menu, Search, Add, or a named item.
class_name / roleElement type.Distinguish inputs, buttons, links, and containers.
boundsPosition and dimensions on the screen.Ground a touch when geometry is needed.
enabled / checked / selectedCurrent control state.Avoid disabled actions and check selection changes.
actionsOperations advertised by the node.Determine whether click, text entry, or scrolling is supported.
Search fields, suggestion rows, content cards, and navigation controls connect to a structured tree. Different properties support different interactions.

Native apps and web pages

Native applications expose interface elements through Android’s accessibility framework. Browsers can expose web roles and links through the same framework. MobiMind processes these two surfaces with separate cleaners and converts their output into a shared semantic representation. The current browser pipeline uses Android accessibility captures; it does not retrieve the page DOM directly.

A conceptual comparison of interface structures. The active browser path uses Android accessibility capture rather than direct DOM retrieval.
04 · Semantic processing

Semantic nodes

Semantic nodes are processed components derived from raw accessibility data. They describe meaningful content and controls using compact labels, roles, state, and actions.

Why process the raw tree?

A raw capture repeats package names, class names, flags, empty fields, and layout information across many nodes. Sending every property to the model increases the input size without necessarily improving target selection.

The cleaner identifies useful elements, consolidates content where appropriate, and removes redundant structure from the model-facing representation. Each retained component receives a semantic ID, such as input_1 or button_1.

Execution data

The server keeps component data, raw node mappings, action mappings, and geometry. The model reads compact text containing semantic IDs and supported actions. It does not need the full execution metadata to select an action.

A semantic ID identifies a target in a particular observation. After navigation or a layout change, the agent must use the IDs from the new capture.

Retained information

Labels, roles, values, useful state, supported actions, and availability indicators help the model identify and operate controls.

Reduced information

Repeated metadata, empty fields, structural wrappers, and duplicate content can be omitted or consolidated in the compact model input.

Processing limitations

Compression can omit useful information. The recorded example below shows the raw capture and the semantic text used in the same session step.

Semantic processing preserves useful meaning and execution mappings. Token figures in the original illustration are illustrative; actual sizes depend on the screen and tokenizer.
Recorded example

Explore a real capture

A Chrome capture of the college faculty page, taken from the session logs. Both representations belong to the same step.

Raw nodes457
Semantic nodes171
Text size reduction97.1%

367,645 raw JSON characters → 10,520 semantic text characters. Raw JSON is measured without formatting whitespace.

Source: 2 October 2026 · bright_wren · step 005.

Token cost

Fewer input tokens can reduce the cost of a model call when the provider charges per token. The recorded example shows the text reduction and a simple token estimate.

Actual tokenization depends on the model, text, and serialization. The comparison excludes the system prompt, task context, history, screenshots, output tokens, caching, and any visual-grounding calls.

Evaluating compression

Measure input size together with target retention and task outcomes. A small representation is useful only if it preserves the information required to select the correct action.

For example, a product choice can be ambiguous if its variant or price is missing. Missing information should trigger another observation, visual inspection, or a user question rather than an unsupported selection.

05 · Incomplete observations

Finding missing targets

The agent must determine whether the target is offscreen, hidden, omitted during processing, or only identifiable from the screenshot.

Missing-target cases and appropriate next steps
CaseExampleNext step
Target is offscreenA product appears below the visible list.Scroll a supported region, capture again, and use the current target ID.
Target is hiddenAlumni is inside a closed navigation drawer.Open Menu and inspect the newly visible links.
Visible control has no useful nodeA map recenter icon is drawn inside a custom view.Describe the target and ground it in the current screenshot.
Node lacks visual meaningGallery images have empty descriptions.Use the screenshot to identify the requested image.
Useful text was omittedA product attribute is present on screen but missing from compact output.Inspect available visual evidence or obtain more information before making a price-dependent choice.
Target is absent or ambiguousNo matching item, or several equally plausible items.Search or navigate further, ask the user, or report the blocker.

Visual grounding

MobiMind can send a target description and the current screenshot to the Moondream grounding component. The returned location supports a tap or another configured gesture.

For “Open the car photo,” the screenshot can distinguish a car from mountains or portraits even when the tree provides no image description. A map recenter control can be located by its icon and position.

Validation before execution

The server rejects a semantic action when its selected target is absent from the current observation. Inventing a semantic ID does not create an executable control.

Visual coordinates also depend on the current layout. After scrolling, opening a dialog, or rotating the device, the agent needs fresh grounding. A successful gesture still requires an outcome check.

Two reasons to use visual evidence: a control missing from the tree, and image content without a useful accessibility description.
Example visual target request · Current step schema
{
  "type": "visual_tap",
  "target": "the map recenter icon at the lower right",
  "tool": "moondream_grounding",
  "action": "tap"
}

The target is a description, not an invented node ID. Grounding supplies the location before the device executes the touch.

06 · Decision and execution

From nodes to actions

The model selects an action and a current semantic target. The server validates that choice and resolves it to executable device information.

Model input

The decision input includes the goal, current semantic screen, relevant memory, session state, and recent actions. An optional image provides visual context.

The output requests one next step. Actions include tapping, entering text, scrolling, changing progress, launching an app, waiting, visual interaction, and asking the user.

Example semantic tap · Step object
{
  "type": "tap",
  "target": {
    "id": "button_1",
    "label": "Menu"
  }
}

This example is valid only when the current observation exposes button_1 with the label Menu. The full decision includes additional session and result fields.

The goal, context, history, semantic screen, and optional image inform one next decision. This illustration includes earlier field syntax; current definitions are in the decision schema.

Targets and supported actions

The server uses the semantic ID to retrieve the raw node and action mapping. Validation checks target existence, applicable state, and action support. A text-entry step requires an appropriate input; a disabled button requires resolving its prerequisite. Android executes the resolved command using the configured mechanism.

Grounding connects a semantic target to a real device interaction. The direct accessibility click shown here is one mechanism; active tap behavior also uses grounded touch and escalation.
07 · Worked examples

Tasks and verification

Each example specifies the request, the relevant observations, the action sequence, and the evidence needed to establish completion.

Open a browser page

Request: “Open the Alumni page.” The initial observation contains Menu but no Alumni link. Tap Menu, capture the drawer, select Alumni, and capture the destination.

Verification: Check the destination heading and available page content. An accepted menu tap is insufficient if the drawer did not open. A loading screen requires waiting and another capture.

Select a product

Request: “Add one 500 ml Amul Milk to the cart.” Identify the search input, enter the query, and inspect the matching product and pack size before selecting Add.

Verification: Check the matching cart item or quantity change. If important attributes are missing from semantic output, inspect the screen before selecting. Several matching variants may require clarification.

Handle a dialog

Request: Delete a selected note. A confirmation dialog displays “Delete note?”, explanatory text, Cancel, and Delete. The agent must resolve the dialog controls from the new capture; background controls are not sufficient.

Verification: After the authorized confirmation, check that the dialog closed and the intended note is absent. Closing the dialog alone does not establish deletion. When user confirmation is required, hold the action until the user responds.

Select an image

Request: “Open the car photo.” A gallery may expose ImageView nodes with empty descriptions. Those labels do not identify which image contains a car.

Verification: Ground the car thumbnail using the screenshot, tap it, and inspect the opened image. If multiple car photos are visible, ask for a distinguishing detail.

Change a setting

Request: “Set the slider to 55.” Read its exposed range and current value. Use a supported progress action, or ground the handle and track if the control requires dragging.

Verification: Read the resulting value. If the setting requires Save or Apply, perform that step only after verifying the adjustment, then check that the value persisted.

08 · Verification

Verification and recovery

Execution feedback reports the command result. Verification evaluates whether the new observation satisfies the expected state change.

Failure and uncertainty

An accepted touch can leave the interface unchanged. Loading, an overlay, a stale target, or incorrect geometry can prevent the intended result.

The agent records the new state and classifies the result as verified, failed, or uncertain. Uncertainty means the evidence is insufficient; it should not be treated as completion.

Recovery actions

Depending on the observation, recovery can include waiting, locating the target again, strengthening grounding, using a screenshot, or choosing a different navigation route.

Repeated attempts use bounded budgets. A retry should use current information and be followed by another outcome check. If execution cannot continue, the agent reports the blocker or requests user input.

Recovery uses the updated screen to re-observe, re-match, strengthen grounding, or change route. The result of a retry still needs to be checked.

Unchanged screen

After tapping Menu, the next capture still shows the closed page. Re-observe the control and retry with a valid target or stronger grounding.

Unexpected destination

The opened page does not match Alumni. Inspect the navigation result and choose an appropriate correction instead of marking the task complete.

Unsaved result

A field displays an edited value, but persistence is required. Save the change and check the stored result before reporting completion.

09 · Context

Memory and context

MobiMind separates reusable knowledge from the temporary state of an active task.

The Android client stores memory in a Markdown file. Relevant sections accompany observations, and the server can return updates for persistence.

A saved navigation note may indicate that Alumni is inside the menu. It helps select a route, but the agent still needs to locate the current control and verify the destination.

GlobalUser preferences

Information that applies across apps, such as preferred language.

PackagesApplication mappings

Friendly application names associated with Android package identifiers.

App scopeApplication knowledge

Navigation details and reusable observations associated with the current package.

Reusable memory is distinct from temporary task state. Current persistence uses global, packages, and app:<package> sections.

Memory versus session state

The current checklist, recent actions, verified facts, and blockers describe the active task. They support the next decision without automatically becoming long-term memory. Browser website notes can be stored in the browser’s app section; separate per-domain retrieval is a possible extension.

10 · User interaction

User interaction

The agent can request missing information, hold an action for confirmation, or stop when the task cannot proceed.

Questions and confirmation

A question resolves information the screen does not provide, such as which of several products the user intended. The answer continues the existing task with its context.

When confirmation is required, the server holds the proposed action. Approval leads to a fresh observation and decision. The user can also stop a running task from the Android interface.

Proposed message signing

The following diagram describes an RSA signing extension for client–server messages. The sender would sign a message with a private key, and the receiver would verify its signature with the corresponding public key.

This diagram is a proposed design. Request and response signing are not implemented in the inspected active communication path.

Proposed communication design. The present-tense wording in the original artwork describes the intended concept, not an implemented signing guarantee.
11 · Research context

Research and evaluation

The main research question is whether a compact semantic representation can reduce model input while preserving the information needed for reliable mobile actions.

Representation quality

Compare raw-node input, semantic input, and semantic input with visual fallback on matched tasks. Measure input tokens, retained target information, target-selection errors, and completed outcomes.

Test difficult cases such as unlabeled icons, hidden controls, repeated products, dialogs, and custom sliders. Missing labels and omitted attributes should be recorded as representation failures when they affect the task.

Execution reliability

Use reproducible initial states and independent final-state checks. Measure task success, false completion, action count, recovery attempts, latency, and provider-reported token usage.

Keep model settings and task conditions consistent when comparing components. The examples on this page demonstrate processing behavior; they do not establish a system-wide success rate or cost reduction.

A conceptual comparison of representations and interaction flows. Broad labels, token estimates, and checkmarks in the illustration are not controlled performance results.
Explore the project

Take a closer look

Replay a recorded run or inspect the screen data behind an agent decision.

Image preview

Real session capture

Raw and semantic nodes

Before and after semantic processing
MetricRaw captureSemantic input
Nodes
Characters
Estimated tokens

smaller text input.

Characters compare minified raw accessibility JSON with the recorded semantic text below. Tokens are estimated as characters ÷ 4, rounded up; they are not provider-reported usage or billing savings.

Raw accessibility JSON

Semantic model input

Layout and repeated metadata are reduced. Labels, roles, supported actions, and availability markers remain in the semantic text. The comparison covers screen data only.