
27 July 2026 · 6 min read
Long-running AI agents: an SME playbook (July 2026)
TL;DR
OpenAI, Anthropic, Microsoft and Google all shipped background-agent capabilities in July 2026. The underlying pattern — an agent working on a queue while a human sleeps, then presenting results for approval — is the same. The winners are SMEs who pick one narrow, boring, repeatable workflow first.
For most of the last two years, an AI conversation ended when you closed the tab. Ask, wait, read, move on. In July 2026 that quietly stopped being true across several of the major vendors — and the shift matters more for a fifteen-person operations team than it does for any one demo reel.
What actually shipped in July 2026
Anthropic shipped Claude Opus 5 on 24 July, and disclosed that the Government of Alberta had already used Claude Code to scan 466 million lines of government code in twenty hours — the kind of run only a long-running agent can make. OpenAI positioned ChatGPT as a partner for extended, multi-step work, and Microsoft 365 Copilot began routing more of that work to the newer OpenAI models. Google added background execution and remote connectors to its managed-agents API — the plumbing that lets an agent pause, wait, and resume against your systems without a human holding the line.
Different vendors, different names. The same pattern underneath.
The vendors caught up to what operations teams already knew: the best AI work is the work you don't have to watch.
What is a long-running AI agent, plainly
A long-running agent is one that keeps working after you close the tab. It picks up a queue of items — emails, invoices, applications, tickets — and works through them step by step for minutes or hours. It calls out to your systems when it needs to. It pauses when it hits something it isn't sure about. And it writes down what it did.
The value is not that it thinks harder. It's that it thinks while you're not there.
The six-step SME playbook
The mistake most companies made in 2025 was starting with the smartest model and looking for a job to give it. The teams shipping value in 2026 do it the other way round. This is the sequence we use with Artellis clients before a single API call is made.
- Pick a boring workflow. Overnight invoice matching. Weekly report drafts. Support-inbox triage. The less exciting it looks on a slide, the higher the odds it ships.
- Write the current process down. Actually write it. Every branch. Every "we usually ask Mary." If you can't describe it, you can't automate it.
- Define done. What counts as a finished item? What has to be true before it lands in a human's inbox in the morning?
- Design the approval step. The human doesn't do the work. The human confirms the work. Two clicks, one row per item, no digging.
- Log everything. Every decision, every source, every tool call. This is your audit trail, your training data and your rollback plan.
- Ship it narrow, then widen. One team, one queue, two weeks. If the confidence and the numbers hold, extend. If they don't, you find out fast.
Where SMEs should be looking, not looking
Look at inbound: quotes, applications, orders, support tickets. Look at reconciliation: invoices versus deliveries, timesheet exceptions, stock discrepancies. Look at drafts: weekly reports, board packs, sales follow-ups.
Don't start with anything that touches money without approval. Don't start with anything a regulator will ask about first. Don't start with the workflow the CEO finds most impressive — start with the one the people who do the work are quietly begging to be rid of.
The trap: model gravity
Every vendor is going to tell you their newest agent runs longer, cheaper and more reliably. In July 2026 that's true for all of them. Which is exactly why the model isn't where your edge is. Your edge is the workflow you pick, the shape of the human checkpoint, and the honesty of the log the agent leaves behind. Swap models later; get those three right first.
Have a workflow you'd like an agent to run overnight?
A discovery week is the shortest honest way to find out whether it's a fit. We pick one workflow, map it, and hand back a ship-or-don't recommendation.
Sources
- Anthropic — Introducing Claude Opus 5 (24 Jul 2026)
- Anthropic — Government of Alberta case study (6 Jul 2026)
- OpenAI and Microsoft 365 Copilot — vendor product pages, July 2026
- Google Cloud / Gemini managed agents — vendor product pages, July 2026
One Workflow
One workflow a week, worth automating.
I look at what your company actually does and write you a short letter: the opportunity, the likely hours it gives back, how involved it is, where the human review sits, and a first step you could run yourself this week. Founding cohort, limited to 100 companies while I personally review every week's recommendations. Reply to any letter and you reach me, not a form.
See a sample letter and how it works
Everything in the letter is decision support: review, test and approve before anything runs in your business. Your email is used for One Workflow only.

