groundy
Developer Tools

Alibaba's Page-Agent: Control Any Website With Natural Language

Alibaba's page-agent embeds an LLM agent in any web page with one script tag for natural language DOM control, with no browser extension or headless browser.

Published Updated 12 references
On this page8 sections

Alibaba’s page-agent is a JavaScript library that embeds an AI agent directly into any web page, enabling natural language control of the DOM without browser extensions, Python scripts, or headless Chrome instances. The repository describes itself as “the GUI agent living in your webpage,” and that is the accurate pitch: one script tag turns an existing web interface into a surface an LLM can operate (GitHub. “alibaba/page-agent.”). It is a thin translation layer between an LLM and the web as it already exists.

What Is Page-Agent?

Page-agent is an open-source JavaScript library published under Alibaba’s GitHub organization that turns any web interface into a surface an LLM can operate. The project is MIT-licensed and available via npm.

Unlike Playwright, which drives a separate browser instance from scripts you host, page-agent embeds directly into the running page via a <script> tag or npm import. The agent lives inside the user’s browser session. It sees the DOM the user sees. It acts with the permissions the user already has.

import { PageAgent } from 'page-agent';
const agent = new PageAgent({
model: 'qwen3.5-plus',
baseURL: 'https://dashscope.aliyuncs.com/compatible-mode/v1',
apiKey: 'YOUR_API_KEY',
language: 'en-US'
});
await agent.execute('Find the highest-priority open ticket and assign it to Alice');

As of early October 2026 the project sits at v1.12.4, capping a rapid run of tagged releases from v1.8.1 through v1.12.4. v1.9.0 added Claude Opus 4.8 support, more robust aborting, and a guard against concurrent execute() calls. v1.10.0 reworked the agent run lifecycle, a breaking change: stop() is now async and run status is decoupled from task outcome. v1.11.0 rewrote the per-model request patching layer and added real-API CI tests for every model on the recommended list. v1.12.0 made extension tab sync stateless so the Chrome MV3 service worker survives idle kills, and v1.12.3 added Kimi K3. (Releases · alibaba/page-agent)

Text-Based by Design

What the repository documents is a modality decision rather than an architecture diagram: the agent manipulates the DOM as text. The README’s feature list is explicit about the contract, no screenshots, no multi-modal LLMs, no special permissions. Everything else follows from that choice.

Text-based browser agents run a loop: serialize the current page into text the model can read, let the model choose the next action, execute it against the live DOM, then re-read the page before the next step. Browser Use works this way when configured for DOM extraction (NxCode), and Gigazine’s hands-on test shows PageAgent behaving the same way: asked for “the list of articles from the day before yesterday,” the agent inspected the page, noticed the search form at the bottom, and used it instead of guessing a date (Gigazine). Re-reading page state at each step is also what lets such agents cope with pages that change under them, including modals, loading spinners, and paginated tables.

The codebase is organized around the components the changelog touches constantly: commit scopes reference core, llms, page-controller, ui, the Chrome extension, and an MCP server the README labels Beta for driving the agent from outside the page. (Releases · alibaba/page-agent)

Notable: the README acknowledges that the project’s DOM processing components and prompt derive from the browser-use project (Copyright 2024 Gregor Zunic, MIT licensed), and states that PageAgent is designed for client-side web enhancement, not server-side automation.

Why the In-Page Position Matters

Every other major web automation approach runs outside the browser:

  • Playwright: drives a separate browser instance from scripts you host, with the infrastructure and credential handling on your side (NxCode)
  • Browser Use: ships Python and TypeScript libraries that run against a local browser, a reused system Chrome profile, or Browser Use’s cloud browsers (browser-use)
  • Computer-use models: Anthropic’s computer-use Claude and OpenAI’s Operator-style agents work from screenshots and can act on any application, not just browsers (NxCode). Anthropic’s Claude for Chrome extension has moved from research preview to beta for Max plan users, which puts it inside the browser session, but it is an end-user extension rather than a developer-embeddable library (Anthropic)

Page-agent’s in-page position eliminates an entire category of operational friction. The agent inherits the authenticated session already open in the user’s browser. There is no separate credential store and no cookie synchronization problem to solve.

Competitors have narrowed this gap from the other side. Browser Use can reuse a system Chrome profile, though its profile sync transfers cookies only, not local storage, IndexedDB, or extensions, and some sites may require signing in again. Stagehand persists cookies in a local profile directory so repeat runs start already signed in. Both still drive a browser from an external process; page-agent runs inside the page itself.

For developers shipping internal tooling (ERP dashboards, customer support interfaces, data entry workflows), this matters. Adding a natural-language command layer to an existing SaaS product becomes a pure frontend problem.

Provider-Agnostic LLM Support

Page-agent ships with no LLM lock-in. The constructor takes a model, baseURL, and apiKey together, and the README’s own example points at DashScope’s OpenAI-compatible endpoint, so any similarly compatible provider is a configuration change rather than a fork. Per-model patches absorb the API variations between providers, from DeepSeek’s tool_choice handling to a transformRequestBody hook for request shaping. The v1.11.0 rewrite of that patching layer came with real-API CI testing for every model on the recommended list, which is the difference between compatible on paper and tested. (Releases · alibaba/page-agent)

Supported as of October 2026:

ProviderIntegration
OpenAI (GPT-5.4 support added in v1.8.1)OpenAI-compatible API
Alibaba Qwen (3.5, 3.6 Max/Flash)DashScope compatible endpoint
Anthropic Claude (Opus 4.8 added in v1.9.0)Per-model API patch
DeepSeek (v4 Flash/Pro)OpenAI-compatible endpoint
Google Gemini, xAI Grok, Kimi (K3 added in v1.12.3), GLMCompatible endpoints
Ollama, LM Studio, LLaMA (local)Local endpoint, no API key

The local options matter as much as the hosted ones: the documentation’s pitch is “fully offline via Ollama” (official documentation), which is the path for enterprises with data sovereignty requirements or air-gapped environments.

Comparison: Page-Agent vs. Competing Approaches

Page-AgentPlaywrightBrowser-UseStagehand
DeploymentIn-page script tag or npm importExternal scripts on your infrastructurePython or TypeScript library, local or cloud browserTypeScript, Python, or Go SDK, local or Browserbase browsers
Interface methodDOM text extractionExplicit selectors you writeDOM extraction, screenshots, or bothAccessibility-tree trimming over Playwright-style APIs
Vision requiredNoNo (no AI at all)OptionalOptional
Session authInherited from the open pageExplicit in your scriptsSystem Chrome profile or cloud profiles, cookies onlyPersistent local profile, repeat runs start signed in
Best forIn-app copilotsDeterministic test suites, high volumeAutonomous multi-step agentsSurgical AI actions, structured extraction
LicenseMITApache 2.0MITMIT

Playwright’s column is the deterministic baseline: an Apache 2.0-licensed test framework with no model in the loop, controlling Chromium, Firefox, and WebKit from five languages (NxCode). The AI columns come from the projects’ own documentation (browser-use, Stagehand).

The closest comparison is Stagehand, which has broadened from a TypeScript library into TypeScript, Python, and Go SDKs running against local or Browserbase-hosted browsers (Stagehand). Its observe() method returns real selectors so credentials never reach the model, a deliberate answer to the same data-exposure question page-agent answers with in-page execution. The remaining difference is deployment: Stagehand is an SDK driving a browser from your code, while page-agent is a script the site itself embeds. Browser Use, meanwhile, has grown a commercial layer on top of its MIT-licensed library: cloud browsers at $0.02 per browser-hour with stealth, CAPTCHA solving, and residential proxies (browser-use).

Installation and Integration

Terminal window
npm install page-agent

For evaluation there is a one-line CDN build pinned to the current version, with the caveat that the demo script routes through the project’s free testing LLM API and is meant for technical evaluation only:

<script src="https://cdn.jsdelivr.net/npm/page-agent@1.12.4/dist/iife/page-agent.demo.js" crossorigin="anonymous"></script>

Appending ?autoInit=false to the script URL loads the library without auto-creating the demo agent, so you can instantiate new window.PageAgent(...) with your own model instead (GitHub. “alibaba/page-agent.”).

A minimal integration for adding a command panel to an existing app:

import { PageAgent } from 'page-agent';
const agent = new PageAgent({
model: 'gpt-5.4',
apiKey: process.env.OPENAI_API_KEY,
language: 'en-US'
});
await agent.execute('Export all overdue invoices from the last 30 days to CSV');

The library also exposes a bookmarklet for quick experimentation, a Chrome extension for multi-tab operation, and a Beta MCP server for agent-to-agent tool calls. (GitHub. “alibaba/page-agent.”, Gigazine’s hands-on test)

The Security Picture

Client-side AI agents that read and act on DOM content introduce a real security surface. The most significant risk is indirect prompt injection: malicious content embedded in a webpage (in a comment, a form value, a dynamically loaded advertisement) that instructs the agent to take unintended actions. Unit 42 documented this moving from proof-of-concept to weaponized, with 22 distinct payload techniques observed in the wild (Unit 42) and attacker intents ranging from SEO poisoning to unauthorized transactions and data destruction.

Because page-agent operates with the user’s authenticated session, a successful injection can reach any resource that user can access: email, file storage, financial systems. This is not a page-agent-specific problem; it applies to every agentic browser tool. But the in-page model concentrates the risk: there is no sandboxed subprocess, no origin boundary, no separate security context.

The project has started narrowing its own blast radius. v1.10.0 disabled the JavaScript execution tool for the multi-page agent and made execute_javascript honor an AbortSignal, so a hijacked run has less room to improvise. (Releases · alibaba/page-agent)

Model-level defenses are improving but not settled. Anthropic’s November 2025 evaluation of its own browser extension put Claude Opus 4.5 at a 1% attack success rate against an internal adaptive attacker (Anthropic), a large improvement over the original launch configuration. The company’s own conclusion is that “no browser agent is immune to prompt injection.”

The free demo version routes data through a server in mainland China, per independent testing (Gigazine); per the developer, the free demo offers only Qwen and DeepSeek, so other models require your own API key. Production deployments should use enterprise LLM endpoints where data residency matters.

The Real Implication: Every Web App Gets an AI Layer

The deeper significance of page-agent is not the technology itself. DOM manipulation via LLMs is a known pattern; the D2Snap research showed downsampled DOM snapshots matching (67% against a 65% grounded-screenshot baseline) or beating (by 8 points in the best evaluated configuration) screenshots on interaction-suggestion tasks as far back as 2025 (arXiv). What stands out is the deployment model. Previous approaches to AI-powered web automation required infrastructure: a Python server, a headless browser farm, credential management, session synchronization. That infrastructure cost made the pattern viable only for well-resourced teams building bespoke automation.

Page-agent’s in-page, npm-installable model changes who can make that trade. The README’s pitch for the SaaS copilot use case is “Ship an AI copilot in your product in lines of code. No backend rewrite” (GitHub. “alibaba/page-agent.”), and the architecture bears the claim out: the agent ships with the page, inherits its session, and needs nothing provisioned beyond an LLM endpoint.

That is the missing layer this tool provides: a thin, LLM-powered translation layer between human intent and the existing web surface, without rebuilding the surface itself.

Whether the broader ecosystem converges on this in-page model or continues to favor external automation frameworks will likely be determined by the security picture as much as the developer experience. A well-publicized prompt injection incident involving an in-page agent with production access could reset adoption curves quickly.

The changelog’s run from v1.8.1 through v1.12.4, spanning new models, a reworked run lifecycle, a stateless extension, and accessibility fixes, says the appetite is there (Releases · alibaba/page-agent). What it adds up to depends on the security answer.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. GitHub. "alibaba/page-agent."github.comAccessed
  2. Releases · alibaba/page-agentgithub.comAccessed
  3. PageAgent.js Official Documentation. "AI-powered GUI Agent."alibaba.github.ioAccessed
  4. browser-usegithub.comAccessed
  5. Stagehandgithub.comAccessed
  6. One Line of Code, Total Web Control — HumanaAI Substackhumanaai.substack.comAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy