Copilot Studio Harness Billing: The Free Pilot Period Just Ended

Finance director reviewing cost dashboards for Copilot Studio agent billing

A finance team at a mid-size distributor spent the summer letting two developers experiment with autonomous agents in Copilot Studio: one to triage vendor invoices, another to reconcile shipping exceptions against purchase orders. Nobody logged the hours. Nobody tracked the credits, because there were none to track. Then, on September 1, 2026, Copilot Studio harness billing changed for exactly this kind of pilot, and both agents started drawing against the tenant’s Copilot Credit pool without a single line of configuration having moved. The invoice for October will look different than the one for September, and the person who approves it may have no idea why.

This is not a hypothetical. Microsoft’s own admin guidance for this change tells organizations to go looking for agents they may have forgotten existed, which is a reasonable ask only if someone already knows to look. For CIOs, CFOs, and IT directors who greenlit a handful of “let’s just see what this can do” pilots earlier in the year, the practical question is no longer whether Copilot Studio’s most capable harness is worth adopting. It is whether the agents already running under it are quietly becoming a budget line nobody planned for.

What Copilot Studio harness billing actually changed on September 1

Copilot Studio ships three distinct harnesses, and the distinction matters more than most licensing conversations acknowledge. A harness, in Microsoft’s own framing, is the runtime layer that sits between an agent’s design and the underlying model: it decides when to call the model, what to hand it, how to interpret what comes back, and which tools to invoke on the agent’s behalf. The Standard harness is deterministic. It follows predefined topics and prompts and behaves the same way every time, which is exactly what you want for a help desk bot answering a known set of questions. The Copilot Chat harness is narrower still, built to ground Microsoft 365 Copilot Chat in an organization’s own content for internal, employee-facing scenarios.

The GitHub Copilot harness is a different animal entirely. It reasons autonomously toward a goal rather than following a script, breaks a task into steps on its own, retries and reroutes when something fails, and can read, edit, and reason over Word, Excel, PowerPoint, and PDF files as part of completing a task. It also supports skills and persistent memory, and runs in a secure sandbox designed for multi-step business processes: an accounts payable agent that reads an invoice, matches it to a purchase order, and routes an exception is the kind of thing Microsoft built this harness to handle. It reached general availability on August 3, 2026, and from that date forward, the billing model was unambiguous: authoring, testing, and evaluating an agent all consume Copilot Credits, not just its published, end-user traffic. The meter starts the moment someone starts building.

What was less obvious is what happened to agents and workflows that existed before that date, built during an earlier preview window when this kind of usage did not consume credits at all. Microsoft’s answer arrived on September 1: those pre-GA agents and workflows, including anything sitting in a Dev or Trial environment, now consume Copilot Credits too, for both the ongoing authoring work a maker does on them and their runtime execution. Standard harness and Copilot Chat harness agents are unaffected. This is specifically about closing the gap between agents that got a free head start during preview and the GA billing model everyone built after August 3 has already been living under.

Why this is a governance problem before it is a cost problem

The credit economics themselves are not exotic once you understand them. Copilot Credits can be drawn from prepaid capacity, priced at roughly $200 per month for 25,000 credits, or purchased pay-as-you-go through an Azure subscription at $0.01 per credit, a rate that runs about 25 percent higher than the prepaid tier. Microsoft’s admin tooling in the Power Platform admin center lets you set a monthly consumption ceiling per agent, with two guardrails: a notification as usage approaches the limit, and a hard stop that disables the agent once it’s reached. Overages beyond an agent’s allocation can either draw from a shared tenant pool or bill against a linked pay-as-you-go plan, and a monitoring dashboard shows month-to-date billed credits per agent alongside a twelve-month trend.

Laptop screen showing an administrative usage and credit consumption dashboard

None of that is the hard part. The hard part is that the September 1 change makes cost visibility retroactive for a category of agent that, by design, was never supposed to need it. A pilot built in July to test whether autonomous reasoning could handle exception routing was, at the time, a genuinely low-stakes experiment. Nobody assigned it an owner with a mandate to monitor consumption, because there was nothing to consume. That agent may still be running today, quietly reprocessing the same exceptions it was built to test, and as of this month it is drawing against the same credit pool as anything built deliberately for production. Multiply that by however many “quick pilot” agents a mid-size enterprise typically accumulates across finance, HR, and operations teams over six months, and the exposure is not a single line item. It is an unknown number of them, some of which may not even have an active business owner anymore.

What to check before the next invoice arrives

Microsoft’s own recommendation is a reasonable starting point: pull historical, non-billed Copilot Credit consumption data from the Power Platform admin center to understand what these agents were actually doing before billing applied to them, and use that history to identify which GitHub Copilot harness agents and workflows exist across every environment, including the Dev and Trial environments that pilots tend to live in and get forgotten in. That inventory step matters more than it sounds like it should, because Dev and Trial environments are exactly where governance attention typically does not reach; they are treated as sandboxes precisely because nothing in them used to cost money.

Once that inventory exists, the decision for each agent is binary and should be made deliberately rather than by default. An agent that proved its value during the pilot period should be formally adopted: given an owner, a consumption ceiling with both guardrails enabled, and a place in whatever ALM process governs production agents. An agent that was built to answer a question that has already been answered, or that nobody has looked at since July, should be retired outright rather than left running under the assumption that “it’s probably fine.” The cost of leaving a stale pilot alive is no longer zero, and the credit allocation and access controls that make sense for a small number of intentional, production-grade autonomous agents do not scale gracefully to an undefined number of forgotten ones.

There is a broader lesson in how this rolled out. Microsoft made the GA billing model clear from day one for anything built after August 3, and it gave organizations a full month, not a surprise cutover, before extending that model backward to preview-era agents. That is a fair way to manage a transition. But it also means the responsibility for knowing which agents exist and what they’re doing sat with each tenant the entire time, and plenty of tenants were not positioned to carry it. As autonomous, credit-metered agents become a bigger part of how Dynamics 365 and Power Platform environments get extended, the organizations that treat agent inventory and consumption governance as a standing discipline, not a one-time reaction to a billing announcement, are the ones that won’t be surprised by the next change like this one. Routeget Technologies has been walking clients through exactly this kind of Copilot Studio environment audit over the past several weeks, and the pattern holds: the agents nobody remembers building are almost always the ones costing the most to leave running.


#CopilotStudio #AgenticAI #GitHubCopilotHarness #CopilotCredits #ITCostGovernance #EnterpriseAI