Embed an in-page GUI agent with Page Agent
Add a JavaScript GUI agent to a web app so users or agents can complete UI tasks through natural-language commands without a headless browser.
npx skills add agentskillexchange/skills --skill embed-an-in-page-gui-agent-with-page-agent
npm install page-agent and instantiate new PageAgent({ model, baseURL, apiKey, language }), or load the documented IIFE script for evaluation. Then call agent.execute('...') from the page context.Use Page Agent when the workflow is inside a web product or browser tab and the agent needs to operate visible UI controls from text instructions. The operator embeds the PageAgent script or npm package, configures the model endpoint, and lets the agent inspect text-based DOM state to click, fill, navigate, or trigger multi-page browser-extension/MCP control when needed. This is not a generic browser automation listing: the boundary is in-page GUI control for product copilots, smart form filling, accessibility commands, and supervised browser-tab workflows where screenshot-heavy or server-side automation would be the wrong tool.
What this skill actually does
Inputs and prerequisites: JavaScript, npm or script tag, LLM API endpoint, optional Chrome extension or MCP-compatible client for multi-page control.
Setup notes: Install with `npm install page-agent` and instantiate `new PageAgent({ model, baseURL, apiKey, language })`, or load the documented IIFE script for evaluation. Then call `agent.execute(‘…’)` from the page context.
Source and verification boundary: use https://alibaba.github.io/page-agent/docs/introduction/overview as the canonical reference before running the workflow; keep commands, API calls, CLI usage, and generated outputs reviewable against that upstream source.
Framework fit: publish this as a Multi-Framework workflow only when the operator can invoke the documented toolchain directly, rather than treating the upstream project as a generic product listing.