Three tiers. You pick.
Everyday drafting uses the Apple foundation model at the core of Apple Intelligence — on your iPhone, built in, instant, even offline. With iOS 27, Apple's Private Cloud Compute steps in when a turn needs more room than the phone has: the same Apple Intelligence stack, on Apple silicon servers, with a much larger context window and deeper reasoning. And when the work is genuinely long and multi-step, Claude takes over on prepaid AI Credit or your own Anthropic key. You switch in the chat itself, or per feature in Settings.
It's an everyday reply, you want it instantly, you're offline, or nothing should leave your phone. This is the default — most messages never need more.
The thread is longer than the phone can hold at once, or you want your Knowledge Base documents searched — still Apple Intelligence, still no key, just more room to think.
The task has steps, you need research with sources, a document produced, or a nuanced tone for a message that matters.
| Capability | Apple Intelligence | Claude |
|---|---|---|
| Where it runs | Two tiers, both Apple. On-device, powered by the foundation model at the core of Apple Intelligence — no network required, works offline. Or Apple's Private Cloud Compute on iOS 27, the same stack on Apple silicon servers, for the turns that need more room. You pick which in the chat | Anthropic cloud — with prepaid AI Credit, or direct to Anthropic on your own API key |
| What you pay | Built in — no API key needed. The Free plan runs unlimited for your first week, then 5 Apple Intelligence actions a day (shared across Smart Reply, Analyze, Summary, Skills, and AI Writer); every paid plan makes them unlimited | Two ways in: AI Credit — prepaid top-ups, no subscription and no API key needed — or the Pro plan plus your own key, paying Anthropic directly per use with no markup |
| Models | Apple Foundation Models — the ~3B on-device model, or Apple's larger server-side model on Private Cloud Compute | Opus 5.5 (most capable for ambitious work), Sonnet 5.5 (most efficient for everyday tasks), Haiku 4.5 (fastest for quick answers) |
| Context window | 4,096 tokens on device (input + output combined, ~3,000 words); about 32,000 on Private Cloud Compute. Longer threads are condensed automatically either way | Up to 1,000,000 tokens (Opus/Sonnet) or 200,000 (Haiku) — handles entire conversation histories. The same on AI Credit and on your own key |
| Input types | Text only — images, PDFs, and documents are handled by Claude | Multimodal — understand and analyse a wide range of visual formats, including photos, charts, graphs, and technical diagrams |
| Extended thinking | Not on the on-device model, which is tuned for direct, single-pass responses. Private Cloud Compute brings the deeper reasoning the larger model allows | Enhanced reasoning capabilities for complex tasks — Claude's step-by-step thought process is visible before it delivers its final answer, with adaptive thinking and adjustable effort (Low, Medium, High, Extra, Max) |
| Reasoning and logic | On device, best for simple focused tasks like everyday drafting rather than heavy logic or math. Private Cloud Compute takes the heavier ones — more context held at once, and reasoning to go with it | Strong complex reasoning, multi-step problem solving, code generation, and mathematical tasks |
| Tool use | Nearly all of them, running through Apple's own frameworks — web search, URL fetching, contacts, messages and chat discovery, calendar, reminders, location, files, and image generation. Knowledge Base document search needs Private Cloud Compute; on the on-device model only saved links can be opened. The Claude-native ones — Task Tracking, Memory, Sub-agents, Code Execution, and Agent Skills — sit outside the Apple set | The full set. Claude decides when to call a tool from your request and the tool's description, then AI Reply Assistant executes it on your device. It adds the agent-only tools on top of everything the on-device set already has. Files and Agent Skills file creation (PPTX, XLSX, PDF, DOCX) need your own Anthropic key |
| Agentic behaviour | Single-turn — one prompt, one response | Claude dynamically directs its own processes and tool usage, maintaining control over how it accomplishes tasks — planning, calling tools, observing results, and refining across multiple iterations |
| Document creation | Not available | Pre-built Agent Skills let Claude run code in a sandboxed container to create PowerPoint presentations, Excel spreadsheets with charts, Word documents, and PDF reports — downloaded and delivered as shareable attachments. Requires your own Anthropic key |
| Message drafting | Generates up to 3 drafts in different tones — you pick the one to use before anything is sent | Human-in-the-loop — Claude drafts messages as cards, you review, choose recipients, and confirm before anything is sent |
| Context management | When a thread outgrows the context window, the transcript is automatically condensed and the conversation continues — no error, nothing lost | Compaction extends effective context length by automatically summarising older context, keeping the active context focused and performant. Prompt caching reduces processing time and costs for repetitive tasks. On-demand tool loading and real-time token tracking |
| Languages | 9 languages (English, French, German, Italian, Portuguese, Spanish, Japanese, Korean, Chinese) | Broad multilingual support across dozens of languages |
| Factual knowledge | A compact on-device model — tuned for everyday drafting rather than broad factual recall or current events | Broader world knowledge with more current training data — stronger on factual and domain-specific questions |
| Privacy | Nothing leaves your phone — maximum privacy | You control when Claude is used. On your own key, drafts go directly to Anthropic; with AI Credit they pass through Flowbie's relay, which stores nothing. |
| Best for | Everyday replies, quick suggestions, full privacy, offline use | Even more expertise — longer drafts, nuanced tone, complex reasoning, multi-step tasks, document creation, and research |
Which Claude, specifically
All three, on either route into Claude — models are never gated by plan or by how you pay, and the context column holds either way, so AI Credit runs every model at its full native window exactly as your own key does. Switch mid-conversation and the next response uses your selection. All three support Extended Thinking. We move to Anthropic’s newest release in each tier as it ships, so these are the current models rather than a fixed list.
| Model | Strength | Context | Best for |
|---|---|---|---|
| Haiku 4.5 | Fastest, most cost-efficient | 200K | Quick drafting |
| Sonnet 5.5 | Best balance of speed + intelligence | 1M | Everyday and complex tasks |
| Opus 5.5 | Most capable for deep reasoning | 1M | Complex, long-running work |
Switching is one setting away
Settings → Chat → AI Features. Smart Reply and Skill each pick Apple Intelligence or Cloud AI independently, and a model row chooses the Claude model used for cloud drafts. Cloud AI only becomes selectable once Claude can serve a request — either a prepaid AI Credit balance or your own Anthropic key — and any feature switches back to Apple Intelligence anytime. On the Apple side there is nothing to configure: a turn runs on Private Cloud Compute when it can serve and on device when it cannot, and in the agent chat you can pin the tier yourself from the model picker.
Writing for each engine
The table above is what each one can do. This is what that means for the instructions you write in a Skill: the same prompt does not get the best out of both.
| When you write | Apple Intelligence | Claude |
|---|---|---|
| Prompt length | Keep it short — 2–4 sentences. The on-device model has a 4,096-token context window (input + output combined), so every word counts. | Can handle longer, more detailed prompts — up to 1M tokens on Sonnet 5.5 and Opus 5.5, 200K on Haiku 4.5, whichever way you pay. Add examples, edge cases, and nuanced instructions. |
| Examples | Use 1–2 simple examples. Too many or too complex will consume the limited context and cause parroting. | Use 3–5 diverse examples for best results. Claude generalises well from richer input. |
| Conditional logic | Keep conditionals simple. The on-device model handles if-this-then-that branching poorly — reach for the tone sliders instead of writing rules. | Can handle moderate conditional complexity — "If the message is a complaint, be empathetic; if it's a question, be direct." |
| Reasoning | Don't ask it to think step by step. At this size the reasoning tends to leak into the reply itself, and the logic is unreliable anyway. | Enable Extended Thinking for complex tasks. Claude handles chain-of-thought reasoning natively with visible, adjustable thinking. |
| Persona depth | A single sentence for role + persona is usually enough. Longer persona prompts eat into the 4K token budget. | You can build richer personas with backstory, communication preferences, and relationship dynamics. |
| Images and documents | Not supported — the on-device model is text-only. No images, PDFs, or file attachments. | Include screenshots, photos, charts, and PDFs. Claude analyses visual content alongside text for richer context. |
Which entry uses which is set per Skill, so a quick everyday Prompt and a research-heavy Skill can run on different engines — templates and the mistakes to avoid are on the Skills page.
What it costs
Apple Intelligence runs on every plan with no per-message fees — the Free plan runs unlimited for your first week, then 5 Apple AI actions a day (Smart Reply, Analyze, Summary, Skills, and AI Writer), and every paid plan makes them unlimited. That covers both Apple tiers: Private Cloud Compute costs nothing extra here, and is metered by Apple against your own device's allowance rather than by us — if you reach it, the app says so and keeps working on device. Claude comes two ways. AI Credit works on any plan: buy a prepaid top-up (from US$4.49), sign in with Apple, and spend it as you go — no subscription, no API key, and it never expires. Or subscribe — Solo US$19.99/year or US$2.99 monthly, Duo US$32.99/year or US$4.99 monthly, Pro US$49.99/year or US$7.49 monthly, each with a 7-day free trial — and bring your own Anthropic API key: you pay Anthropic's standard usage rates directly with no markup, can cap spending in your Anthropic console, and add Managed Agents, Agent Skills, and Files. Bringing a key works on all three plans; what they differ on is how many messaging accounts you connect at once. Context is the same either way — each model runs at its full native window, so Sonnet 5.5 and Opus 5.5 give you 1M tokens whether you paid with credit or a key. Subscribing also puts AI Credit in your balance once your first bill goes through — US$1.00 on Solo, US$1.50 on Duo, US$2.00 on Pro, once per Apple Account — so you can try Claude without buying a top-up. Either way, a typical drafted reply uses only a few thousand tokens, and you choose the model per task — Haiku for fast, inexpensive drafting up to Opus for the hardest work.
Your personal AI assistant for messaging.
Download AI Reply Assistant on iOS and try it free.
Apple Intelligence
Claude