AI agents: what can they actually do?
From a good answer to a finished task. A practical look at the capabilities, the limits, and the human decisions in between.
Not just
about technology.
About people, too.
Independent
AI magazine
Since 2026
From chat replies to actions in the real world.
From a good answer to a finished task. A practical look at the capabilities, the limits, and the human decisions in between.
Background notes preserve conventions, decisions, and durable facts.
An OpenAI case study looks at testing, software changes, and production monitoring.
The managed service brings the Codex harness to longer-running sessions.
Google introduces a dynamic alternative to fixed-rate video processing.
The retrieval layer adds a loop for navigating and checking documents.
Cohere introduces orchestration for multi-step enterprise processes.
Google integrates interface interaction into its main Flash model.
Anthropic starts with Slack channels and work assigned through mentions.
Google DeepMind describes layered oversight for increasingly capable internal agents.
Medium 3.5 arrives with remote tasks that can continue beyond an open terminal.
The orchestration layer targets durable, observable AI processes.
OpenAI introduces cloud agents for shared, multi-step work.
Mistral engineers describe a specialized loop for generating and checking tests.
The team joins work on Claude’s ability to understand and operate interfaces.
OpenAI’s macOS interface centers on longer tasks and parallel development.
Mistral adds configurable subagents, skills, and clarification to its terminal workflow.