The CLI is now as important as the UI

I have been building Microsoft Foundry agents with GitHub Copilot as my coding partner. The biggest lesson is simple: for agentic software, the command-line interface is becoming as important as the user interface.
For years, good software experience mostly meant a good user interface. That still matters. But agent experience matters too, and agents work best with a command-line interface backed by a proper API. As more tedious process work moves to agents, the CLI becomes commercially important: it affects cost, scalability, security, observability and ultimately the economics of software.
Agents are not magic. They need surfaces they can trust.
There is a temptation to think computer use and browser use solve everything. If a human can click through a SaaS product, surely an agent can do the same. Well, yes, sort of.
Large language models (LLMs) are fundamentally text-based systems. They can process screenshots and visual layouts, but usually at greater token cost and with less precision than structured text. Browser use may read the structure of a web page without seeing it exactly as a person does and can get blocked by bot protection systems. Computer use goes further, repeatedly taking screenshots, interpreting the screen, clicking, typing and checking the result. Clever, but not cheap or especially clean when done at scale.
Browser and computer use have their place for prototypes, testing and awkward integrations. But for serious workflows they are a poor default: less efficient, harder to secure, harder to observe, harder to make deterministic and harder to guardrail than a well-designed API or CLI.
Flat per seat AI subscriptions also hide the problem. Frontier-lab computer use can feel free at the point of use while quietly brute-forcing its way through screenshots using expensive capability. Fine for a prototype or one time use, but suboptimal for scaling cost effectively.
The CLI is the better path for agents
A better pattern is boring: give the agent a proper CLI, backed by a well-designed API. The agent can inspect commands, read responses, update code, rerun workflows and compare results. That creates the engineering loop software needs: small changes, repeatable tests, visible failures, cleaner logs, narrower permissions and safer failure modes.
It also makes iteration quicker. The agent can inspect the command, read the response, update the code, rerun the workflow and compare the result. That is how software improves: small changes, repeatable tests and visible failures.
Cost control, observability and security all get easier too. CLI-driven workflows produce cleaner logs, clearer telemetry and more predictable consumption. They can run with narrow permissions, in controlled environments, with auditable output and safer failure modes.
Code-first has appealing advantages
Because LLMs are text-based and current AI systems are particularly good at coding, a CLI shifts the centre of gravity away from clicking around a UI and towards building in code. Commands, configuration files, deployment scripts, prompts and tool definitions all fit naturally in a Git repository, where humans and agents can inspect changes, understand the system, roll back mistakes and compare environments.
Considering a lot of UI driven technology change is never documented, the self-documenting nature of code is very appealing. In a code-first workflow, the agent definition, infrastructure, permissions and deployment process can all be version controlled. You can review changes, roll back mistakes and compare environments.
The same agents that help you build the system can also explain it later. Six months from now, when you have forgotten why a tool was scoped a particular way or why a deployment script has an odd exception, a coding agent can inspect the repository and reconstruct the logic. It cannot do that reliably if the important decisions were never documented after the configuration was done in the portal manually.
Model fungibility is not academic
That leads to the next question: what should you build on? The earlier points only really pay off if you still have room to choose the right tools underneath the agent, especially the model.
Plenty of companies have gone straight to frontier labs because it is the fastest way to get something working. That is understandable, but it comes with a catch: you are often building inside that lab’s models, pricing, tooling and product assumptions. Those models may be the most capable answer, but they are also often the most expensive one. More importantly, the harness can start to shape the architecture before you have made a conscious decision about lock-in.
Model fungibility matters because the model, pricing, tooling and product assumptions you start with can quietly shape the architecture. A model-agnostic platform gives you a better chance of choosing the right model for the task. You may still standardise on one frontier lab for good commercial reasons, but that should be a deliberate choice, not a trap you discover after agents are embedded.
Observability is the difference between iteration and superstition
Your agent probably will not work properly first time. It may look good at first and then randomly fail on the third run. That’s normal. Agents, like humans, become good through iteration, not first attempt perfection.
Observability is what turns agent development from hopeful fiddling into an engineering loop. Your coding agent needs to see what the model saw, which tools it called, what each tool returned, what failed, what retried and what it cost. Without that evidence, developers fall back on prompt tweaks, vague tool changes and model swaps because something felt better in testing.
Without traces, developers drift into superstition: prompt tweaks, vague tool changes and model swaps because something “felt better” in testing. The trace is what turns agent development from hopeful fiddling into an engineering loop.
Your agent activity data is an asset
Agent activity data is not exhaust. It is the raw material for optimisation: failed tool calls, corrected outputs, human approvals, rejected recommendations and successful completions all show how work happens. Govern it properly with retention rules, access controls and redaction, but do not casually delete the learned intellectual property of how your agent performs your work.
Consumer agents are not business ready
Proper enterprise agents can feel slow and fiddly compared with consumer bots, but that friction is the cost of security, observability and accountability. Consumer agents can rely on broad context, vague permissions and patchy auditability because the user is usually accepting the risk personally. That is not good enough for work a business must own.
A business agent needs to be as professional as any human performing a business process – it needs to be able to articulate what it did and why and be able to improve and learn. If it can’t do that, it has limited business value and is just a risky gimmick.
Control the infrastructure if you want to control the economics
If you want to control agent economics, control enough of the infrastructure to see what is happening. In Azure, that means separating environments, monitoring token and tool costs, and understanding the workflows that drive consumption before they become another mysterious cloud bill.
Agent workloads are not just application hosting with a nicer chat box: they can involve model calls, retrieval, tool execution, tracing, storage, queues, orchestration and human approval, each with its own cost profile.
This can sound horribly technical for a small business, but coding agents make the learning curve unusually accessible; the limiting factor is often time to tinker, not the ability to understand the stack.
Bottom line
My advice is simple: buy market-leading SaaS tools for core business processes where they expose proper APIs and CLIs, then build agents and niche apps around them. Build code-first, keep model choices flexible, insist on traceability, govern activity data and understand the economics. The UI still matters, but if the API and CLI are second-class citizens, your agents will be too.
If you’re trying to make sure your team are well setup to take advantage of the benefits of AI, we can help. Reach out below.👇