Getting started with AI — Layer 3: handing over the keyboard — getting AI to do things for you

This is the third in the series. In Layer 1 we had a conversation. In Layer 2 we got good at feeding it context, and I signed off promising to connect AI to your own documents and tools. That's still coming — but while I was drafting it, the ground moved. The most useful version of "connect it to your stuff" turned out not to be uploading files at all. It's this: the AI can now use your stuff the same way you do — through a browser, logged in. So we're doing that first.

Everything in Layers 1 and 2 had one shape: you ask, it answers, you go and do the thing. Draft the email — you send it. Compare the options — you fill in the booking form. The AI was a brilliant colleague who could never leave the meeting room.

That's the wall that just came down. The one idea in this layer:

AI has moved from answering to doing. You can now hand it a small task — not a question, a task — and it will open a browser, click through the site, fill in the forms and bring you back the result.

The industry word for this is an "agent", which sounds grander than it is. An agent is just an AI that can use a computer the way you do: look at the page, decide what to click, type into boxes, move to the next page. Nothing mystical — it's the same assistant from Layer 1, given hands.

Why this only became useful in the last few weeks

Here's the boring, decisive fact that held all of this back: almost everything worth doing online is behind a login. Your orders, your bookings, your invoices, your council account, your dentist's appointment page. An assistant that gets stopped at every "Sign in" screen can only do toy tasks on public pages — which is why the early agent demos felt like party tricks.

The login problem was genuinely hard to solve safely. Nobody sensible wants to paste their passwords into a chat window, and no AI company sensibly wants to hold them.

That's what changed, and it changed twice in a fortnight. As I write, in late August 2026:

  • OpenAI (ChatGPT) added secure sign-in to its cloud browser on 25 August — the AI works on a remote browser, and when it hits a login page it stops and hands you a secure form. You type your password and any two-factor code yourself. The AI never sees them; they're not stored; the signed-in session then persists for future tasks until it expires.
  • Anthropic (Claude) made its Chrome extension generally available on 26 August — it takes the opposite route: Claude works inside your own browser, using the sessions you're already logged into. It never needs your password at all, because you logged in yourself, the normal way, possibly months ago.

Same wall, two different doors. Both sensible. Which one suits you depends on the task, so let's take them one at a time.

Route 1: ChatGPT's cloud browser — a task you can walk away from

This lives in ChatGPT's paid plans (Plus and Pro) under ChatGPT Work, the multi-step "go and do this" side of the product. When a task needs the web, it runs a browser on OpenAI's computers, not yours — which means the task keeps running after you shut the laptop.

What it's actually like to use:

  1. You describe the task like you'd brief a person: "Find me three quotes for hire of a 3.5t van, one day, collection in Teddington, two weeks Saturday. Fill in the quote forms with my details below. Stop before anything is paid."
  2. It opens the sites and works through them. You can watch it move, or leave.
  3. When a site wants a login, it stops and shows you a secure sign-in form — you check the web address it's about to sign into, enter your password and any two-factor code, and it carries on. Your password goes straight to the remote browser; the AI model itself never sees it, and it isn't stored. Password managers work in the form.
  4. Anything that commits you — a booking, a payment — is put in front of you to confirm, one at a time. It doesn't spend your money on its own judgement.

One thing to know because it's both convenient and worth managing: after you've signed into a site once, that signed-in session sticks around for future tasks until it expires. Handy for the sites you'll use weekly; if you'd rather a site was forgotten, you clear it yourself under Settings → Cloud browser → Browser data.

Best for: tasks with waiting in them, multi-site legwork, anything you want to hand over entirely — renewals, appointment hunting, gathering quotes, cancelling the thing that takes four screens to cancel.

Route 2: Claude in Chrome — a colleague at your elbow

Claude's version (paid plans, rolling out across tiers as I write — install from the Chrome Web Store, or start at claude.ai/chrome) sits in a side panel of the browser you already use. There is no separate login story to learn, because it's your browser: whatever you're signed into, it can work with, and your passwords never enter the picture.

Using it feels less like dispatching a task and more like handing the keyboard to someone sitting next to you:

  1. Open the site you want help with, open the Claude panel, and say what you want done: "Go through this month's orders and copy the invoice totals into the spreadsheet in my other tab."
  2. It clicks, types and fills forms on the page in front of you, and you can see every move.
  3. It asks permission site by site — you grant access per site, and can revoke it in settings — and it checks with you before the sensitive moments: entering personal details, downloading files, granting access to something new. Some things it simply won't do at all, even if you ask — purchases and other financial transactions are blocked outright, so anything involving money stays with you.

Best for: fiddly work inside sites you're already living in — forms, admin screens, moving information between tabs, the "twenty minutes of clicking" class of job.

What to actually try first

Concrete beats clever here. Good first tasks share three traits: they're bounded (a clear end), checkable (you can see whether it was done right), and low-stakes (a mistake costs you a shrug). For instance:

  • fill in a quote or enquiry form on three suppliers' sites with the same details;
  • book the dentist/MOT/table from the choices you've already made;
  • cancel or reschedule the thing whose cancellation page is deliberately buried;
  • gather your last six months of invoices from a supplier portal into one folder;
  • fill in a repetitive government or council form you already know the answers to.

Notice what's not on the list: nothing open-ended ("sort out my admin"), nothing high-stakes. That's not because the tools can't be pointed at bigger things — it's because Layer 1's rule still applies: you get a feel for a tool by watching it do small real jobs, and correcting it, before you trust it with consequential ones.

The sensible-caution section

Same deal as Layer 2's honesty about made-up facts — here's where the edges are:

  • Watch the first few runs. Not because disaster is likely, but because that's how you learn what it's good at, exactly as you would with a new assistant.
  • Keep it away from banking, legal and medical accounts for now. Both companies effectively say the same, and block or guard the riskiest categories. The convenience isn't worth being an early adopter there.
  • Confirm-before-commit is your friend — leave it on. The good defaults keep every commitment with you: ChatGPT puts each booking or payment in front of you to confirm, and Claude won't make purchases at all. Don't loosen any of that to save taps.
  • Only give it tasks on sites you trust. A malicious page can try to smuggle instructions to an AI reading it ("prompt injection", if you want the term to search). The vendors are actively hardening against this and the numbers are improving — it's the main reason to keep early tasks boring and low-stakes.
  • The login designs are genuinely careful — that part I'd trust: ChatGPT's form keeps your password away from the model entirely; Claude never asks for one. The thing to stay thoughtful about isn't the password handling, it's what you authorise the AI to do once inside.

That's Layer 3

One idea again: answering has become doing. The assistant you've been talking to since Layer 1 can now open a browser, get past the login — safely, at last — and finish the small task itself. Start with jobs that are bounded, checkable and low-stakes, watch the first few runs, and you'll quickly develop the only instinct that matters: knowing which jobs to hand over.

And next time, the promise from Layer 2, properly: connecting AI to your own documents and tools — your files, your records, your systems — so it doesn't just do tasks for you, it does them already knowing your world.


I'm Simon — I run Simon Studios, an AI-first development studio in Teddington. I help small businesses and individuals work out where AI genuinely helps (and where it doesn't), then build it — no jargon, no hype. Everything in this post is the personal, off-the-shelf edge of what I build: when a repeated job matters enough, I build a dedicated agent around it — connected to your own data and tools, hosted and kept working for you. If that sounds useful, that's the conversation I love to have. Book a free intro call at cal.com/simonstudios/free-intro-call and we'll work out whether there's something real here for you.