Two years ago most enterprise AI projects were chat interfaces bolted onto a knowledge base. In 2026 the conversation has moved on. Models are reliable enough at tool use, structured output and multi-step reasoning that teams now ask a different question: can the system do the work, not just talk about it? That is what people mean by agentic AI, a model that plans, calls APIs, reads the results and decides what to do next until a goal is met.
The upside is real, but so is the risk. An agent that can issue a refund, update a CRM record or merge a pull request can also do those things wrongly, at machine speed. This article summarises how RixlSoft approaches agentic systems for clients: which use cases are ready, what the architecture looks like, which guardrails matter, how to evaluate behaviour before and after launch, and where to start without betting the business.
Where agents are ready today
The best early candidates share three traits: the work is high volume, the steps are describable, and a mistake is detectable and reversible. That rules out a lot of glamorous ideas and rules in a lot of unglamorous ones. Back-office operations are the sweet spot, because the process already exists in someone's head or a runbook, and the agent is mostly orchestrating systems that have APIs.
We steer clients away from fully open-ended agents that are told to go and improve revenue. Narrow agents with a clear definition of done outperform general ones, are cheaper to run, and are far easier to evaluate.
- Customer service triage: classify, pull order and account context, draft a resolution, and execute low-risk actions such as resending an invoice.
- Finance operations: invoice matching, exception handling in accounts payable, and reconciliation notes for human review.
- Sales operations: lead research, CRM hygiene, meeting preparation and follow-up drafting.
- Engineering: dependency upgrades, test generation, incident summarisation and first-pass code review.
- IT and HR service desks: access requests, onboarding checklists and policy questions with ticket updates.
A reference architecture that holds up
Under the hood, a production agent is less magical than the demos suggest. It is a loop: the model receives a goal, context and a list of tools; it proposes an action; your code validates and executes that action; the result goes back to the model; repeat until done or until a limit is hit. Almost all of the engineering effort sits outside the model, in the orchestration layer.
We build that layer with a few fixed components. A tool registry exposes typed functions with strict input schemas, increasingly through the Model Context Protocol so the same tools can be reused across models and clients. A retrieval layer supplies grounded context from documents and systems of record. A state store keeps the plan, intermediate results and a full trace of every step. A policy engine decides, per tool and per argument, whether an action runs automatically, needs approval, or is blocked. Finally, an observability pipeline records prompts, tool calls, latencies and costs so every run can be replayed.
Keep the model swappable. Pricing and capability shift every few months, and the teams that hard-code one provider's quirks into business logic pay for it later. A thin abstraction over model calls, plus an evaluation suite, lets you change models with evidence rather than hope.
Guardrails that actually reduce risk
Guardrails are not a single content filter at the end. They are layered controls, and the most effective ones are ordinary software engineering. The single most important rule: the agent should never hold more permission than the human it is acting for, and ideally much less.
Prompt injection deserves special attention. Any text the agent reads, an email, a web page, a support ticket, can contain instructions. Treat all retrieved content as untrusted data, keep high-impact tools behind explicit approval, and never let the output of one untrusted source directly trigger an irreversible action. The OWASP guidance for LLM applications is a useful checklist here.
- Least-privilege credentials scoped per tool, with read-only access by default.
- Hard limits on steps, spend, records touched and monetary value per run.
- Schema validation on every tool call, rejecting anything outside allowed values.
- Allow-lists for destinations such as email domains, payment accounts and repositories.
- Idempotent actions and a documented rollback path for anything that writes data.
Evaluation before and after launch
If you cannot measure an agent, you cannot improve it or defend it. Before launch we build an evaluation set from real historical cases, typically a few hundred, each with an expected outcome. We score the final result, but also the path: did it call the right tools, in a sensible order, without unnecessary steps? Deterministic checks cover what they can, such as correct record IDs or totals, and model-graded rubrics cover the rest, with humans spot-checking the graders.
After launch, evaluation becomes monitoring. Sample live runs for human review, track task success, escalation rate, cost per task and time to completion, and feed every failure back into the test set. Any change to prompts, tools or model version should run the full suite in CI before it ships, exactly like a code change.
Human-in-the-loop as a dial, not a switch
Human oversight works best when it is designed as a graduated dial. In the first phase the agent drafts and a person approves everything. As accuracy is proven on a category of task, that category moves to approve-by-exception, where only low-confidence or high-value cases go to a reviewer. Only well-understood, reversible actions ever become fully autonomous.
Design the review experience carefully. A reviewer should see the agent's proposed action, the evidence it used and its reasoning summary on one screen, and be able to approve, edit or reject in seconds. Those decisions are also your best training and evaluation data. For regulated contexts, including systems that may fall under the EU AI Act's obligations as they phase in, this audit trail is not optional.
Cost, and where to start
Agent costs are driven by tokens per step multiplied by steps per task. Long contexts and chatty loops add up quickly. Practical levers include routing simple steps to smaller models, caching stable context such as system prompts and policies, summarising history instead of replaying it, and capping steps. Always compare cost per completed task with the fully loaded cost of the human process, not cost per call.
Our recommended starting point is a single process, one team and a six to ten week pilot. Pick a workflow with clear success criteria, build the evaluation set first, launch in draft-only mode, and expand autonomy as the numbers justify it. That path produces a working system and, just as importantly, the internal confidence to build the next one.
Resist the urge to build a platform before you have a use case. Shared components such as the tool registry, policy engine and tracing pipeline are worth standardising, but only after one or two agents have shown which abstractions you actually need. The second agent should reuse far more than the first.
Key takeaways
- Choose narrow, high-volume, reversible workflows for your first agents.
- Most of the engineering lives in orchestration, tools, policy and observability, not the model.
- Build the evaluation set before the agent, and run it on every change.
- Treat human oversight as a dial that opens only as measured accuracy improves.