본문 바로가기

전체 글133

How Mobile World Model Guides GUI Agents? How Mobile World Model Guides GUI Agents?Recent advances in vision-language models have enabled mobile GUI agents to perceive visual interfaces and execute user instructions, but reliable prediction of action consequences remains critical for long-horizon and high-risk interactions. Existing mobiarxiv.org 2026. 5. 19.
Agent+P: Guiding UI Agents via Symbolic Planning Agent+P: Guiding UI Agents via Symbolic PlanningLarge Language Model (LLM)-based UI agents show great promise for UI automation but often hallucinate in long-horizon tasks due to their lack of understanding of the global UI transition structure. To address this, we introduce AGENT+P, a novel framework tarxiv.org 2026. 5. 19.
Agentic Reward Modeling: Verifying GUI Agent via Online Proactive Interaction Agentic Reward Modeling: Verifying GUI Agent via Online Proactive InteractionReinforcement learning with verifiable rewards (RLVR) is pivotal for the continuous evolution of GUI agents, yet existing evaluation paradigms face significant limitations. Rule-based methods suffer from poor scalability and cannot handle open-ended tasks,arxiv.org 2026. 3. 24.
OpenClaw-RL: Train Any Agent Simply by Talking OpenClaw-RL: Train Any Agent Simply by TalkingEvery agent interaction generates a next-state signal, namely the user reply, tool output, terminal or GUI state change that follows each action, yet no existing agentic RL system recovers it as a live, online learning source. We present OpenClaw-RL, a fraarxiv.org 2026. 3. 17.