AI Agents vs Agentic AI: A Builder’s Taxonomy Based on What Actually Runs

Share





FuturPulse analysis · 11 October 2026

A 2025 review by Ranjan Sapkota, Konstantinos Roumeliotis and Manoj Karkee separates task-specific AI agents from agentic AI systems built around multi-agent collaboration, dynamic task decomposition, persistent memory and coordinated autonomy. Talorys is a concrete AI-agent example: its published design describes one personal agent with tools, state and scheduled reminders. Agent Memory Repo offers a concrete agentic-system ingredient: its swarm example describes several agents sharing findings, questions and a common memory repository.

At a glance

  • The taxonomy paper defines AI agents as modular systems for task-specific automation and describes agentic AI through collaboration, decomposition, persistent memory and coordinated autonomy.
  • Talorys documents a single-user personal agent with SQLite-backed durable state, tool calling and scheduled alarms.
  • Agent Memory Repo documents a clone, search, update and push memory loop that persists knowledge between sessions.
  • Strands Box documents policy enforcement outside an agent process for shell, Python, network and MCP operations.
  • The useful procurement question is not “is this agentic?” but which system properties are documented, inspectable and enforceable.

What is the practical difference?

An AI agent performs a bounded job; agentic AI describes a wider system that organises work over time. The distinction is architectural, not a property permanently attached to one language model. The same model can answer a prompt, call a tool or participate in a system that delegates work among several actors.

Sapkota, Roumeliotis and Karkee describe AI agents as modular systems enabled by language models for task-specific automation. Their abstract contrasts them with agentic AI systems marked by multi-agent collaboration, dynamic task decomposition, persistent memory and coordinated autonomy. That definition is useful because it describes observable system design rather than marketing language.

An AI agent might retrieve a customer record, validate a form or draft a reply. It receives an input, chooses a next action within its permissions and produces an output. Its work can involve several tool calls without becoming a wider agentic system.

Agentic AI is the arrangement around those workers. It can break a goal into work packages, retain shared context, decide who handles each package and govern how actions proceed. That does not mean every agentic system needs every possible component, but it does require more evidence than a chat interface with tools.

Moveworks makes a similar enterprise distinction: an individual agent handles a defined task, while agentic AI coordinates agents, data sources and tools across broader workflows. That is a vendor explanation, not a standard. Still, it points builders towards the right evidence: plans, hand-offs, records and permission boundaries.

A practical test follows. Call a product an AI agent when one software actor acts toward a bounded objective using approved tools. Call the deployment agentic AI only when its surrounding system manages work, state and constraints across steps or actors.

Which four runtime properties should builders test?

Test for decomposition, durable state, coordination and external control. The first three align closely with the characteristics named in the taxonomy paper’s abstract. External control is a separate FuturPulse audit category because a system needs limits that do not rely on the model following instructions.

  • Decomposition: the system turns a broad objective into smaller tasks, assigns or sequences them, and can revise that plan after receiving new evidence.
  • Durable state: facts, task status or decisions survive a model call and remain available for later inspection, correction or deletion.
  • Coordination: agents or services use explicit roles, shared records and a process for handling conflicting work.
  • External control: software outside the model decides which files, code, credentials, network requests or tool calls are allowed.

These are independent properties, not stages on a ladder. A useful single agent can have durable memory but no coordination. A multi-agent workflow can coordinate tasks yet lack a policy layer that restricts external actions.

Decomposition is more than a list of steps in a prompt. The system should hold a plan somewhere, show what remains and update it when evidence changes. Otherwise, the apparent plan may be only a transient model response.

Durable state is also more than the model’s context window. Context is information supplied for one interaction. Durable state survives that interaction and has an owner, a storage location and a way to be amended.

Coordination requires more than several agents accessing the same folder. Shared storage can support coordination, but it does not by itself assign responsibilities or settle contradictions. Builders should ask who owns a decision when two workers produce incompatible updates.

External control means that an agent cannot simply ignore its own restrictions. A model can be instructed not to send a request or read a file. A separate enforcement process can deny the request even if the model attempts it.

Freshworks describes AI agents as task-focused programs and agentic AI as more autonomous, adaptive goal pursuit. Autonomy alone, however, is not enough for this audit. A timer can start one agent without a user prompt, while the system still lacks decomposition, coordination and controls.

What do the repositories actually document?

The reviewed repositories document separate building blocks: personal state, shared memory and policy enforcement. They should not be treated as interchangeable evidence of agentic AI. Each addresses a different operational problem.

What does Talorys demonstrate?

Talorys describes itself as one personal AI agent running inside the user’s Cloudflare account. Its architecture names a TalorysAgent Durable Object with SQLite storage for conversations, memories, tasks, notes, projects, automations, sessions, settings and usage. The same published architecture names Workers AI for streaming and tool calling.

The repository says alarms run reminders and recurring schedules. It also says memories can be viewed, edited and deleted, which is useful evidence that the project’s personal facts are intended to be managed rather than merely appended. Those details support the durable-state point in the audit table.

Scheduled action is not one of this article’s four properties. It shows that the application can act after a timer fires, but it does not demonstrate task decomposition or coordination between multiple agents. The repository’s “single-user by design” description points in the opposite direction: one owner and one named personal agent.

Talorys also documents adjustable limits for output tokens, context tokens, tool calls, reasoning steps, AI requests and scheduled AI runs. Those are application guardrails. The published description does not establish an external policy engine that independently approves or denies every action, so they are not scored as external control here.

What does Agent Memory Repo demonstrate?

Agent Memory Repo specifies a memory loop in which each session clones the latest memory, searches it, updates entries and pushes after every edit. The project says the memory lives in a separate repository and can persist across sessions when suitable persistent storage is available. That supports a durable-state score.

The repository also describes using the same memory repository for agents working in parallel. It says each agent pushes changes and Git flags conflicting edits. This is useful infrastructure for shared work, but it is not enough to award a coordination point under the definition above.

The project’s swarm example goes further than simple shared memory. It describes agents assigned to database, runtime, cache and load-balancer areas, plus common findings and questions files. It also instructs workers to pull and read those records before each step.

That example shows how a user could build coordinated work around the memory format. It does not document a built-in orchestrator that assigns roles, resolves every conflict or enforces a shared workflow. The table therefore scores the published memory component, not the fuller swarm scenario described around it.

Agent Memory Repo explicitly says automatic startup and scheduled Dreaming are not included in its local trial. Its documentation describes Dreaming as a dedicated agent that can periodically add and clean up memory. That feature description does not alter the score because the audit tests documented runtime properties of the repository’s core memory design.

What does Strands Box demonstrate?

Strands Box documents an enforcement process outside the agent sandbox. Its host operating system enforces direct access restrictions, while its shell, Python interpreter, egress gateway and MCP broker send handled operations to the Dogwood policy engine. This supports the external-control score.

Box says checked operations are denied by default unless a matching permit rule allows them. It further says a forbid rule overrides a permission. That arrangement matters because the agent does not make the final policy decision.

The project documents a shared event history for policies based on earlier actions and elapsed time. For example, a file read through shell or Python can affect whether a later outbound HTTP request is permitted. This is policy state, not evidence of an agent planner or a multi-agent coordinator.

Box also says its gateway can authenticate permitted requests without giving the agent the underlying credentials. It records policy decisions in OTLP JSON by default. Those records can help an operator inspect what the enforcement layer allowed or denied.

How do the repositories score on the test?

FuturPulse counts only the four documented runtime properties defined above. One point is awarded where a project’s published material clearly supports one property. The count is an architectural reading, not a benchmark, security certification or reliability rating.

Builder audit: documented agentic-system signals
ProjectSignals documentedWhat the published design demonstratesWhat it does not establish
Talorys1 of 4Durable state in a SQLite-backed Durable Object.Dynamic decomposition, multi-agent coordination or external policy enforcement.
Agent Memory Repo1 of 4Durable memory stored in a separate Git repository across sessions.Built-in role assignment, an orchestrator or complete conflict resolution.
Strands Box1 of 4External policy control over selected operations and credentials.Planning, durable task memory or a complete agent application.

The audit applies the paper’s collaboration, decomposition and memory criteria alongside the published designs for Talorys, Agent Memory Repo and Strands Box. A score of 1 does not mean the projects are equivalent; it means each clearly documents one property in this narrow test.

Each reviewed repository documents one audited system property: Talorys, Agent Memory Repo, Strands Box
Each reviewed repository documents one audited system property · Source: doi.org

The table is deliberately strict. Talorys has alarms, but scheduled action is not a fifth scoring category. Agent Memory Repo can be used in a swarm, but shared memory alone does not prove documented coordination.

Strands Box can restrict an agent’s actions, but it does not claim to decide a business objective. Combining these components can produce a more complete system. The components should still be assessed separately before a buyer accepts an “agentic” label.

What are the five common types of AI agents?

The usual categories are simple reflex, model-based, goal-based, utility-based and learning agents. They classify how an individual agent chooses an action. They do not directly classify whether a deployed system is agentic.

A simple reflex agent responds to an immediate condition. A model-based agent maintains an internal representation of its environment. A goal-based agent compares available actions against a desired outcome.

A utility-based agent weighs outcomes using defined preferences. A learning agent changes its behaviour using feedback or new data. A single workflow can combine these decision methods without gaining persistent shared memory or multi-agent coordination.

Moveworks lists reactive, model-based, utility-based and learning agents in its enterprise explanation. Its examples place these agents within defined operational roles. That fits the distinction here: decision style describes a worker, while agentic architecture describes the system connecting workers.

Builders should avoid treating these labels as a maturity ranking. A reflex-style approval check can be preferable when the rules are fixed and easily audited. A learning component can create a governance burden when changed behaviour is hard to explain.

Is ChatGPT generative or agentic AI?

ChatGPT is generative AI when it produces content from a prompt. It can become part of an agent when surrounding software provides tools, state and rules for taking action. A chat response alone does not establish an agentic system.

The taxonomy paper positions generative AI as a foundation for agents, which add tool integration, prompt engineering and reasoning enhancements. It then describes agentic AI as the broader arrangement involving collaboration, decomposition, persistent memory and coordinated autonomy. The important unit of analysis is therefore the deployed system, not the model name.

A plain content-generation product displays an answer. An agent can use an approved tool after producing a plan. An agentic system can retain task state, route work and record how an action was approved, denied or revised.

That distinction is especially important when a product advertises “agentic” chat. Ask what happens after the browser session ends. Ask where the task state lives, what can trigger a later action and which component can stop a risky request.

Where do hosting and deployment fit?

Hosting makes an application available; it does not make the application agentic. Hugging Face says Spaces can host machine-learning demo applications using Gradio, Docker or static HTML. Its documentation also says a Space can be upgraded to accelerated hardware.

Those are deployment choices. They do not reveal whether an application has durable memory, decomposes work or independently enforces permissions. A polished hosted demo can conceal important runtime details.

Hugging Face describes Spaces as a way to host demo apps on a user or organisation profile. That makes them useful for showing an interface. Buyers should still inspect what happens after a session closes and what permissions the backend holds.

Good architecture descriptions use operational language. “One agent with three tools and a daily trigger” is clearer than “agentic platform.” “Several workers share a repository, while a coordinator assigns tasks” is clearer when that is the actual design.

Who is making each claim, and why does that matter?

The academic review and the repository documentation support the strongest architectural claims here. The Sapkota paper is listed as an Information Fusion journal article and was last revised on 30 September 2025. It offers a conceptual taxonomy, not a universal compliance test.

Talorys publishes its personal-agent architecture and MIT licence in its repository. Agent Memory Repo publishes a memory format, installation instructions and usage examples. Strands Box publishes an Apache 2.0-licensed sandbox engine and its enforcement design.

Vendor explainers can clarify how companies use the terms in practice. They also promote product categories. That is why the audit gives priority to repository architecture and the paper’s stated definitions, rather than accepting a category label as evidence.

What could we not verify?

No source reviewed here establishes a universal threshold at which an AI agent becomes agentic AI. The academic taxonomy, enterprise explanations and open-source projects use overlapping but not identical language. The four-property test is a transparent editorial framework, not an industry standard.

Repository descriptions do not establish real-world reliability, cost, incident rates, planning quality or human-override frequency. Those claims would require repeatable evaluations, task definitions, logs and published measurements.

Agent Memory Repo’s swarm scenario illustrates coordination practices, but the repository documentation does not establish that its core memory component automatically orchestrates agents. Talorys documents a personal agent with durable state and schedules, but not multi-agent work allocation. Strands Box documents enforcement, but not an end-to-end agent application.

The next useful evidence would be a reproducible system specification showing task assignment, stored state, approval points, conflict handling and action logs. Until then, buyers should ask vendors and maintainers to describe what runs rather than rely on the word “agentic.”

Which sources support this analysis?



Maya Chen
Maya Chen
Maya Chen covers AI agents, orchestration frameworks, tool-use, and evaluation. She focuses on what actually works in production—failure modes, safety boundaries, and measurable performance—without the hype.

Read more

Local News