Today, when an AI agent calls our render_pdf tool with delivery: "url", it gets back a signed link. The agent pastes that link into the chat, you click it, a new tab opens, and the conversation you were having is now behind you. MCP Apps is the extension that closes that gap: an MCP server can return an interactive HTML interface that renders inside the conversation. This article is a design exploration, not a release note. PDF4.dev does not ship an MCP App today.
What is MCP Apps?
MCP Apps is an official extension to the Model Context Protocol, identified as io.modelcontextprotocol/ui. It lets a server attach an interactive HTML interface to a tool, which the host renders inline in the chat. The official overview describes the motivation plainly: "Text responses can only go so far. Sometimes users need to interact with data, not just read about it."
The mechanism combines two primitives that already exist in core MCP. A tool declares a pointer to a UI resource in its metadata, and the server serves that resource as an HTML document. The overview lists four steps: the host can preload the UI from _meta.ui.resourceUri before the tool is even called, it fetches the resource, it renders the HTML in a sandboxed iframe, and then the app and host exchange JSON-RPC messages over postMessage.
The extension is deliberately narrow. Per the specification, the initial revision supports HTML only, served with the MIME type text/html;profile=mcp-app, under the ui:// URI scheme.
Is MCP Apps stable or experimental?
It is official, not experimental. The MCP blog post of 26 January 2026 announces MCP Apps as "the first official MCP extension" and describes it as ready for production. The extensions overview reserves the word "experimental" for repositories carrying the experimental-ext- prefix, and ext-apps is not one of them.
That said, two caveats are worth stating. The build guide closes with a note that "MCP Apps is under active development" and points readers at the GitHub issue tracker. And the extensions overview says extensions "evolve independently of the core protocol", with updates managed by the extension repository maintainers rather than core maintainers. So the identifier is stable, and breaking changes are supposed to arrive as a new identifier such as a -v2 suffix, but the surface around it is still moving.
How does a host negotiate MCP Apps?
Through the standard extensions capability mechanism, in both directions, and always opt-in. The extensions overview is explicit: "Extensions are always disabled by default and require explicit opt-in from the developer."
A client advertises support inside _meta["io.modelcontextprotocol/clientCapabilities"] on each request, declaring the extension identifier and the MIME types it can render. A server advertises its side in the server/discover response, under capabilities.extensions. That second half matters for anyone who read our piece on the 2026-07-28 stateless spec: with the initialize handshake gone, server/discover is where capability declaration now lives.
| Piece | Value |
|---|---|
| Extension identifier | io.modelcontextprotocol/ui |
| Client declaration | _meta["io.modelcontextprotocol/clientCapabilities"].extensions |
| Server declaration | server/discover result, capabilities.extensions |
| UI resource scheme | ui:// |
| UI resource MIME type | text/html;profile=mcp-app |
| Tool metadata pointer | _meta.ui.resourceUri |
| Transport between app and host | postMessage, JSON-RPC |
Graceful degradation is part of the design. The overview recommends that "a server offering UI-enhanced tools should still return meaningful text content for clients that don't support the UI extension." For a PDF tool, that rule is not a nice-to-have: the base64 payload and the signed URL have to keep working untouched for every client that never sees the iframe.
What does the server code actually look like?
Two registrations and one metadata field. The official build guide publishes this pattern, using helpers from the @modelcontextprotocol/ext-apps package:
import {
registerAppTool,
registerAppResource,
RESOURCE_MIME_TYPE,
} from "@modelcontextprotocol/ext-apps/server";
const resourceUri = "ui://get-time/mcp-app.html";
registerAppTool(
server,
"get-time",
{
title: "Get Time",
description: "Returns the current server time.",
inputSchema: {},
_meta: { ui: { resourceUri } },
},
async () => ({ content: [{ type: "text", text: new Date().toISOString() }] }),
);
registerAppResource(
server,
resourceUri,
resourceUri,
{ mimeType: RESOURCE_MIME_TYPE },
async () => ({
contents: [{ uri: resourceUri, mimeType: RESOURCE_MIME_TYPE, text: html }],
}),
);On the browser side, the guide uses an App class from the same package, with app.connect() to establish the channel, an app.ontoolresult callback for the result the host pushes in, and app.callServerTool() for calls the UI initiates on its own. The guide is careful to note that the class "is a convenience wrapper, not a requirement": the postMessage protocol is documented and you can implement it directly. Starter templates exist for React, Vue, Svelte, Preact, Solid and vanilla JavaScript.
What can a sandboxed app do, and what can it not do?
The security model is the reason a host is willing to render third-party HTML at all. The specification requires views to render in sandboxed iframes with restricted permissions, requires hosts to enforce CSP based on declared domains, and requires all communication to flow through auditable JSON-RPC messages.
| The app can | The app cannot |
|---|---|
| Call tools on its own server through the host | Access the parent window's DOM |
| Receive tool inputs, partial inputs and results pushed by the host | Read the host's cookies or local storage |
| Request a display mode change and report size changes | Navigate the parent page |
| Push structured data back into the model's context | Run scripts in the parent context |
| Request open-link, camera, microphone, geolocation or clipboard-write permissions | Reach domains it did not declare in its CSP metadata |
The default CSP, when a resource declares no metadata, is deny-by-default: default-src 'none' with connect-src 'none', allowing only self-hosted and inline scripts and styles. An app that needs external fonts, a CDN or an API call has to declare those origins in _meta.ui.csp, which exposes connectDomains, resourceDomains, frameDomains and baseUriDomains. The host is also free to refuse capabilities: the overview notes a host "might restrict which tools an app can call or disable the sendOpenLink capability."
What PDF4.dev's render_pdf returns today
Our MCP render_pdf tool has exactly two delivery modes, and both end in a dead end for the eye. Reading the handler:
delivery: "base64"(the default) returnspdf_base64plussize_bytesandduration_msin the structured result. The agent holds a binary blob it cannot look at, and a 2 MB PDF becomes roughly 2.7 MB of base64 in the context window.delivery: "url"writes the bytes to disk, mints an HMAC-signed token, and returns a URL valid for 24 hours. The context cost drops to a few hundred bytes. The viewing cost moves to you, in a browser tab.
Both are correct API design. Neither lets the model or the user see the document without leaving the conversation. The agent renders an invoice, reports "done, 3 pages, 184 KB", and has no idea that the footer overlapped the total on page 2.
What an interactive preview would change
A UI resource for a render result would turn a link into a viewer. The obvious build is the one the official pdf-server example already demonstrates in the ext-apps repository: an interactive PDF viewer built on PDF.js, with tools for navigation, search and annotation, plus an app-only tool that streams the file bytes in chunks with pagination metadata.
| Moment | Today | With a UI resource |
|---|---|---|
| Seeing the result | Open a signed URL in a new tab | Page through the PDF inside the chat |
| Fixing a layout bug | Describe the problem back to the agent from memory | Point at the page, then ask for the change |
| Iterating on a template | Re-render, re-open the tab, compare two tabs | Re-render into the same view |
| Changing page size | New tool call, new URL, new tab | A control in the app calls render_pdf again |
| Context cost | Small with URL delivery, large with base64 | Small: the bytes go to the iframe, not the transcript |
The last row is the one I keep coming back to. app.callServerTool() means the UI can re-render without the model spending a turn on it. A preset switcher, a margin slider, a sample-data field: those are interface concerns, and routing them through the language model is both slower and more expensive than a button. Meanwhile ui/update-model-context gives the app a way to tell the model what the user just did, so the conversation does not lose the thread.
The read-only case is just as useful. Our preview_template tool returns an image today, which is fine for one page and useless for a twelve-page contract.
Why we are not shipping it yet
Three reasons, in order of weight.
First, delivery. An MCP App is a bundled HTML document served as a resource, and the guide's recommended path bundles the UI into a single file with vite-plugin-singlefile. Our MCP server is a route handler inside a Next.js app, not a standalone server with a build step for a second frontend. That is solvable, not free.
Second, the bytes. The pdf-server example streams a file through an app-only tool in chunks. We already store renders on disk behind a signed URL with a 24 hour TTL, so the app could fetch that URL directly, but only if the host's CSP allows connectDomains to reach our origin. That is a real constraint to design against, not a detail.
Third, honesty about reach. The client matrix is community-maintained, and the checkmark on a client tells you the extension is implemented, not that the specific behavior we need works the way we expect. Every UI-enhanced tool still has to return a correct text and structured result for hosts that never render the iframe. Until we have tested that fallback properly, shipping the iframe would be shipping a promise.
What is still undocumented
Some things I wanted to state and could not verify from a primary source, so I am not stating them. The pages I read do not give per-client limits: nothing on maximum resource size, iframe height, how long a rendered app stays alive in a conversation, or whether a view survives a page reload. The client matrix marks support as a binary checkmark with no version or feature granularity, and it carries an explicit note that it is community-maintained. If you are building against a specific host, test that host.
Where this leaves us
MCP Apps closes a gap that our own tool surface makes obvious every day: we can generate a document in under a second and then hand the user a link. The extension is official, the identifier is stable, the security model is restrictive in the right direction, and the reference implementation for a PDF viewer already exists in the ext-apps repository. What is missing on our side is a bundling story and a tested fallback, not a spec.
If you want the shape of our current MCP server first, we wrote up how we built it: 14 tools, 4 resources, 3 prompts, and the single source of truth that keeps the docs honest. That is what any UI resource would sit on top of.
Sources: the MCP Apps overview, the build guide, the extensions overview, the client matrix, the 2026-01-26 announcement, and the extension specification.
Free tools mentioned:
Frequently asked questions
- What is MCP Apps?
- MCP Apps is an official extension to the Model Context Protocol, identified as io.modelcontextprotocol/ui. It lets an MCP server return an interactive HTML interface that the host renders inline in the conversation, instead of only text, images or structured data.
- Is MCP Apps experimental or stable?
- It is an official extension, not an experimental one. The MCP announcement of 26 January 2026 calls it "the first official MCP extension" and describes it as ready for production. The build guide still says MCP Apps is under active development, so treat the surface as official but still moving.
- Which clients support MCP Apps today?
- The official client matrix on modelcontextprotocol.io marks MCP Apps as supported by Claude (web), Claude Desktop, VS Code GitHub Copilot, Microsoft 365 Copilot, Goose, Postman, MCPJam, ChatGPT, Cursor, Archestra.AI and PostHog Code. That matrix is community-maintained, so verify against the client you actually target.
- How is a UI resource declared?
- The server registers a resource under the ui:// scheme with the MIME type text/html;profile=mcp-app, then points a tool at it through the _meta.ui.resourceUri field. The host fetches the resource and renders it in a sandboxed iframe.
- Can an MCP App read my Claude conversation or my cookies?
- No. The spec requires the view to render in a sandboxed iframe. It cannot access the parent DOM, read the host's cookies or local storage, navigate the parent page, or run scripts in the parent context. Everything goes through an auditable postMessage JSON-RPC channel that the host polices.
- Does PDF4.dev ship an MCP App today?
- No. Our render_pdf tool returns either a base64 PDF or a 24 hour signed URL. This article is a design exploration of what an MCP App preview would add, not a release note.
- What happens to a UI-enabled tool on a client that does not support the extension?
- Extensions are opt-in on both sides and degrade gracefully. The extensions overview recommends that a server offering UI-enhanced tools still return meaningful text content for clients that do not support the UI extension, which is exactly how a PDF tool should behave.
Start generating PDFs
Build PDF templates with a visual editor. Render them via API from any language in ~300ms.



