CopilotOnToast: Making GitHub Copilot CLI a Little More Noticeable
A small PowerShell utility that adds Windows toast notifications to GitHub Copilot CLI hook …
When organisations discuss the risks of generative AI, the conversation usually settles into familiar territory. What will it cost? Where will our data go? Which models are approved? Who owns the subscription? These are necessary questions, but they are no longer enough for the tools that we are now placing into the hands of software engineers, citizen developers, and people who simply want to build something.
I was reminded of this while reading an article, “Token bleed: Why AI consumption velocity is your next chief operational risk”. The article’s central warning is a good one, that token consumption creates financial exposure at machine speed, while the financial controls used by many organisations still operate retrospectively, a month in arrears. By the time the invoice has arrived and travelled through the reporting cycle, the money has already been spent.
That argument reflects risks I have seen discussed in real organisations, and with the clients I work with. It also prompted a wider thought: the token bill is only the most visible consequence of poorly governed AI tooling. It has a unit price, appears on an invoice and can eventually be attributed to a team, a tool or a user. The cost of a compromised dependency, an exposed secret, an incorrectly executed database command, or sensitive information being sent somewhere it should never have gone is much harder to predict. It can range from an embarrassing clean-up exercise to lost clients, reputational damage, regulatory penalties (GDPR fines anyone?), or something severe enough to threaten the organisation itself.
The conversations I have been involved in around approving AI Coding Assistants have reinforced this concern. In large organisations, the executives who ultimately own the financial, technology, or cyber risk are rarely the people reviewing an individual request. A local budget holder or operational leader may be treated as the “authoriser”, even when the decision being placed in front of them is fundamentally technical or commercial in nature. Their approval can then depend heavily on the person asking for the tool, who may be enthusiastic about getting access but may not understand, or may unintentionally understate, the risk being accepted.
This is not just about what appears on the rate card or subscription page of your AI tool of choice. It is about what the tool can reach, what it can change and whose authority it is using while it does so.
This is a long article because the risks are connected and they deserve consideration, but you do not need to read it from beginning to end if you don’t want to. If one area matters more to you, jump straight to it:
Part of the problem is that we continue to use broad terms such as “AI assistant” or “copilot” for tools with very different abilities. A chat assistant is primarily an ask-and-respond interface; it can write code, but that often means an isolated script, a simple one-file program, or an answer that the user must manually transfer into another environment. When the code does not work, the conversation can circle around repeated revisions without the model ever having a reliable view of the wider system.
An in-editor copilot moves beyond that. It has access to the file being edited and often to other parts of the codebase. It can use the context of the project, understand existing patterns, and propose changes that fit the work already under way. The user remains close to the editing process, but the assistant is no longer responding only to text pasted into a chat box.
A full coding agent is different again; tools in this category can inspect repositories, edit files across the repository, invoke terminal commands, run builds, install packages, execute tests, use authenticated command-line tools, interact with source control, and connect to other services (the list goes on). Increasingly, they can also access the user’s browser, call APIs, and use tools exposed by MCP servers so they can operate beyond the development environment. In practical terms, a full coding agent can attempt many of the actions available to the human engineer sitting at the machine, only faster, across more tasks and do it without getting tired.
That capability is what makes these tools transformative. It is also why treating a coding agent as if it carries the same risk as a chat window is a serious and fundamental error.
The most useful mental model I have found is to treat a full AI coding assistant like a very junior member of the engineering team. Imagine someone with enormous energy, a desire to impress, and access to a vast amount of book knowledge; but without the experience that turns knowledge into professional judgement. They do not understand the history of your organisation, its unwritten conventions, its risk appetite, its production incidents or why a particular policy exists. Now imagine an entire fleet of those junior engineers working at machine speed; then place them under the supervision of someone who is still developing the judgement needed to challenge them”
No responsible engineering leader would give that fleet unrestricted access to every system and accept every change without supervision. Yet that is close to the operating model being created when a powerful coding agent is installed on a corporate laptop, given broad local permissions and placed in the hands of someone who cannot evaluate its actions.
This is where I find the distinction between a vibe coder and an AI-accelerated software engineer useful. I am not using “vibe coder” as a dismissive label for every non-engineer who creates something with AI. It describes an operating mode: accepting code, commands, and architectural choices because the output appears to work; without understanding how it works or what consequences follow from it. Of course a professional engineer can also fall into that behaviour; we have all done it at some point. The difference is not the job title on the person’s email signature, but whether informed engineering judgement remains part of the process.
The relationship is not linear. Capability and reach define the potential blast radius, while autonomy determines how quickly the agent can act within it. As that autonomy increases, the risk can grow much faster than the user’s ability to recognise and contain a bad decision.
Risk pressure ∝ Agent capability × Reach × (Autonomy ÷ Informed oversight)
AI did not create every risk discussed in this article. Insecure dependencies, misplaced secrets, destructive commands, and poor data handling all existed before generative AI and they continue even when it isn’t used. What AI changes is the friction. It allows a weak decision to be repeated, expanded, and executed at a scale that previously required considerably more time or expertise. Those without knowledge and experience would often stop when it got complex or confusing; now they can just keep saying “try again”, burning more credits but also forcing it to “just work” without consideration of the “how” or “why”.

Token economics remains part of this story because poor engineering decomposition can translate directly into unnecessary spend. Consider a user who wants an AI coding assistant to create a processor for a large CSV or collection of Excel files. The objective may be to calculate totals, group records, produce summaries or transform the data into another format. To write that program, the assistant usually needs to understand the columns, data types, representative values, constraints, and expected output. It rarely needs every row of the production dataset.
A disciplined approach would inspect the file structure first. For a CSV, that could be as simple as opening it in Excel, reviewing the columns and identifying their types. The user could then give the assistant the schema, a small safely constructed sample and a clear description of the required processing. Once written and tested, the resulting program can run conventionally against the real source file. At that point the bulk operation is ordinary compute, not an ongoing conversation with a language model.
Instead, I have seen and heard examples ranging from CSVs of considerable size to discussions about multi-gigabyte collections of files being handed wholesale to an assistant in order to write processing code. In some cases the program is eventually run outside the assistant anyway. In others, the user simply follows with “now run it”, asking an agentic tool to process the entire dataset when deterministic code could have performed the task at a much lower cost and more predictably.
Different products handle large attachments in different ways. Some may extract, sample, index or divide the information rather than inserting every byte into a single model prompt. The precise token charge will therefore vary. The defensible principle is not that every uploaded byte always becomes a billable token. It is that indiscriminate ingestion creates avoidable AI processing and makes it harder to understand both the cost and the data exposure.
Cost is only one concern. A large file may contain commercially sensitive information, personal data, client material or information subject to far stricter controls. Depending on the product, contract and configuration, content may be retained for telemetry, made available for support and abuse monitoring, or treated differently between consumer and enterprise offerings. In the most sensitive organisations, the consequences of putting the wrong dataset into the wrong tool extend beyond embarrassment or regulatory cost; they may even reach into government and critical-infrastructure contexts. What is your organisation’s blast radius?
The engineering lesson is not simply “write shorter prompts”. It is to decide what the model genuinely needs in order to reason, and to leave deterministic bulk processing to ordinary software. Token literacy can be directly connected to engineering literacy. The person who understands the problem can separate the small reasoning task from the large execution task. The person who does not may pay an AI model to repeatedly consume information it never needed. One example I managed to step in front of for a client resulted in some back-of-a-napkin maths that would have cost upwards of $20,000 in token costs for a single run based on the size of the input files, thank goodness they already had a limit on file upload sizes.
The same lack of friction appears when an agent decides that the solution to a problem is another dependency. This is normal engineering behaviour. Mature third-party packages allow teams to reuse well-understood capabilities rather than rewriting them badly. The concern is not that an AI assistant uses open source software; it is the speed and basis on which it may choose what to bring into the organisation.
An agent may find a package that appears to match the task, add it to the project and continue until the build passes. It does not automatically inherit the organisation’s risk appetite, approved technology choices or collective memory. It may not know that a once-popular package has fallen out of favour, that maintainers have warned users away from it, or that a recently disclosed compromise occurred after the model’s knowledge was assembled. It may choose an obscure dependency to save a few lines of code without judging whether the additional attack surface is justified.
An experienced engineer can still get this wrong. Professional experience is not a magical shield against a convincing package, a compromised maintainer or an attack hidden deep in a transitive dependency tree. It does, however, make certain questions more likely.
A vibe coder may see only that the confusing red error message has disappeared.
The threat is broader than obviously malicious new packages. Typosquatting and dependency confusion can direct a user towards the wrong component. A trusted maintainer’s account can be compromised, allowing an attacker to publish a malicious release of an otherwise mature package. An abandoned project can change ownership. A dependency can introduce a vulnerability, an unacceptable licence or a maintenance burden even when nobody involved is acting maliciously.
Different ecosystems expose these risks in different ways. npm packages can run lifecycle scripts; Python packages can execute build backends and install executable code; modern NuGet PackageReference packages do not simply run an npm-style post-install hook, but can affect builds through MSBuild targets, analyzers, source generators, tools and included binaries. Also use “npm” and “Python” consistently. The technical route varies, but the underlying principle is consistent: retrieving a dependency can bring executable behaviour into a process running with the user’s access level (which could even be that of an administrator).
Professional software delivery usually surrounds this decision with additional controls. Trusted registries, lock files, software composition analysis, dependency review, source and application security testing, peer review and controlled CI/CD pipelines each contribute levels of protection. None is perfect, but together they reduce the chance that one weak decision moves silently into production. A citizen developer working through a conversational tool may not know that these controls exist, let alone know how to implement or respect them.
Even where an organisation has approved registries or package controls, it must confirm that the agent actually uses them. A tool with open internet access may attempt to retrieve a package from a public source because that is the most direct way to complete the task; and when it faces a blocker it will often try to work around it. A policy written for human engineers is not automatically a control over agent behaviour, and an agent is rarely concerned about being fired for accessing a questionable website at work.
One of the most highly used npm packages was compromised in March 2026; Axios versions 1.14.1 and 0.30.4 added a hidden dependency called [email protected]. Axios did not use that dependency in its application code. Its purpose was to run a post-install script that downloaded a cross-platform remote-access trojan on Windows, macOS, and Linux. Microsoft noted that ordinary npm install or automated dependency updates were sufficient to trigger it. The GitHub advisory recommended treating affected computers as fully compromised and rotating all accessible secrets that the machines had visibility of. The industry knew about it at lightning speed, nothing talks faster than a group of disgruntled engineers, but LLMs were still suggesting the package to users. Check out the Microsoft Security analysis and GitHub advisory.

The most immediate security issue is also one of the easiest to overlook: a desktop coding assistant usually operates with the user’s local permissions. Sandboxing and approval models vary between products, and many tools include meaningful safeguards. Those safeguards still depend on how the product is configured, what the process can reach and whether the user reviews or bypasses its requests.
When someone enables an unrestricted or “YOLO” mode, the model does not acquire administrator privileges by magic. The user has delegated part of their existing authority without retaining meaningful supervision. That authority may include access to source code, adjacent documents, network shares, authenticated command-line sessions, cloud subscriptions, source-control platforms, internal services reachable through the corporate network or VPN, and environment variables containing API keys or tokens.
Environment variables are not inherently a bad way to pass configuration to a process. The problem is the accumulation of long-lived secrets in a user’s environment, often copied from the quickest “getting started” instructions for a command-line tool. A warning may say to use a proper secret-management system in production, but someone who does not understand the reasoning may follow the first working instructions and never revisit the decision. Months later, credentials unrelated to the current task may still be visible to any process running as that user.
Consider the sorts of actions a capable agent can propose or execute. It can delete a file tree because it concludes that generated files are blocking a build. It can issue a database command, decide that a failed upgrade is easier to solve by dropping and recreating the database, and act on the wrong environment because the user never established proper separation between development and production. It can access a logged-in browser and take an action on a third-party service because the action appears to satisfy the task. The human user may recognise that the browser action is dangerous, but only if the tool pauses and the user understands what they are being asked to approve.
Even a familiar destructive command demonstrates the problem. To an engineer, rm -rf is immediately recognisable and demands attention to its target. To a user who cannot read shell commands, it is just another string of text between the request and the promised result. The confirmation prompt offers little protection if the only available decision is “allow the thing I do not understand” or “stop the tool from completing my work”. For those that don’t know, rm -rf means nuke the current directory.
This connects directly back to the software supply chain. If malicious package-supplied code runs during installation or build, it inherits whatever the process can access. A plausible attack can discover secrets, archive code and configuration, transmit the archive and remove obvious traces. Attacks using these techniques are not hypothetical in the wider software ecosystem, and AI can make the route to executing them faster, more frequent and less scrutinised.
Many organisations respond to these concerns by pointing out that the human remains “in the loop”. The phrase is often used as if the presence of a confirmation dialog settles the risk. It does not.
A meaningful human control requires more than a person being present. The person must understand the proposed action, recognise its likely consequences, have enough context to identify when something is wrong and be willing to stop the process. For a professional engineer reviewing a familiar command, an interruption can be a valuable control. For someone who cannot interpret the command, it is mostly friction. Friction is also exactly what “always allow” and autonomous modes are intended to remove.
This does not mean professional engineers are immune to automation bias, haste or misplaced trust. Experienced people make mistakes, especially when a tool has completed the previous fifty actions correctly and trained them to accept the next one. The difference is that professional engineering provides layers of knowledge and process around the individual decision. Peer review, environment separation, controlled pipelines, automated security testing and operational experience all make it more likely that a bad action will be questioned or contained.
Nor does it mean non-engineers should be excluded from AI-enabled creation. These tools can make experimentation, prototyping and local automation available to people who previously could not build those things for themselves. That is a real benefit. The appropriate capability and environment must, however, match the person’s ability to understand and contain failure.
True experimentation should take place in a sandbox where access to sensitive data, production systems, and unrelated corporate resources is deliberately constrained. Production-grade engineering should retain the protections expected of production-grade software: distinct environments, controlled CI/CD, software composition analysis, static and dynamic testing where appropriate, peer review and accountable ownership. A prototype does not become safe because someone calls it an experiment after connecting it to real data and operational systems.
Training therefore needs to cover more than prompting. Vendor certifications from the likes of GitHub and Anthropic can provide useful product knowledge, but knowing how to get a better answer is not the same as knowing how to supervise execution. Users need to understand permissions, data handling, dependencies, environments, secrets, review, and escalation. Formal training from experienced engineers is sometimes treated as a blocker by people who “just want to build”. In reality, it is part of the safety net that allows them to build without unknowingly accepting risks on behalf of everybody else.

An organisation may know exactly how many AI-assistant licences it owns while knowing far less about what those assistants can do. It may not have a consolidated view of which tools can execute commands, which autonomous modes are enabled, what internal systems they can reach, which package sources they use, where inherited secrets are stored or whether actions can be attributed to a user or agent identity.
The governance decision is too often reduced to “AI access: yes or no”. That makes tools with radically different blast radii look interchangeable. A more useful starting point is to classify the capability being introduced.
| Tool use | Typical capability | Principal concern |
|---|---|---|
| Chat and explanation | Generates text from supplied context | Data handling, accuracy and token cost |
| In-editor assistance | Reads and proposes changes within code context | Code quality, intellectual property, context access and review |
| Agent with command execution | Edits, installs, builds, tests and invokes tools | Local authority, supply chain, secrets and network reach |
| Agent connected to enterprise systems | Acts through APIs, MCP servers or service identities | Delegated authority, auditability, cross-system blast radius and runaway action |
This is not a universal risk rating. Products, configurations and organisational environments differ. It does demonstrate why approving a full coding agent as though it were simply another chat-based SaaS product leaves important questions unanswered.
Risk ownership makes this harder. The CIO, CTO, CISO, CFO or another executive may ultimately be accountable, but an end user asking for a vibe-coding tool will rarely speak to them. The request passes through local leadership and an IT service process, where a budget holder can become the de facto risk approver. Procurement may conduct a SaaS assessment, security may review contractual data handling, and finance may approve the subscription, but people rarely asks what happens when the installed application executes commands as the user and connects to other systems.
The formal process is also easily bypassed. A member of staff hears about the latest tool through a conference, advert or LinkedIn post, is frustrated that procurement has not approved it, pays with a credit card and expenses the subscription. The organisation now has an agentic development tool running on its estate without the enterprise guardrails, contractual position or technical assessment that a centrally managed product might provide.
A serious review needs to ask how the tool integrates with other systems, what additional reach those integrations create, what enterprise controls are available, and whether existing professional engineers consider it suitable for the environment. It should distinguish genuine enterprise capability from exciting product marketing. Most importantly, the people approving the risk must either understand it or have access to people who do.
The answer is not to ban every powerful tool. Their capability is precisely what makes them useful, and removing it entirely would sacrifice much of the engineering benefit. The operating principle should instead be straightforward: increase access and autonomy only alongside visibility, constrained authority and demonstrated competence.
All of the necessary control patterns are achievable today but none are free. Managed sandboxes, package controls, monitoring, training, security assessment and expert review introduce cost and can delay adoption. Those trade-offs should be considered honestly, but inconvenience does not make the underlying risk disappear. A fast approval is not evidence of a safe decision.
Experimentation can be given considerable freedom inside an appropriate sandbox. The environment should be designed so that a mistake cannot reach sensitive commercial or client data, production services or unrelated areas of the corporate estate. Real data should not be required merely to discover whether a new tool or approach has promise. If the concept cannot be tested with synthetic or safely prepared information, it is already moving beyond low-risk experimentation.
Corporate desktops and laptops should themselves be treated as part of the organisation’s operational environment. They contain code, credentials, documents, authenticated sessions and network access that matter to the organisation. Only approved and controlled tools should operate there, particularly when they can execute commands or act through the user’s identity. Approval should consider the complete operating model, not only whether the vendor promises that customer prompts will be handled responsibly.
Practical controls can include monitoring software installations across the estate, blocking download locations for unapproved tools, routing dependencies through trusted registries, managing secrets so they are not broadly inherited by local processes, and correlating token consumption with tool actions and identity access. Training remains essential, but it must develop operational judgement rather than merely teaching users how to obtain more impressive output.
Teams need a safe route to experiment, a clear route to production and access to professional engineers who can help them cross the boundary. Guardrails should raise the floor for less experienced users without lowering the ceiling for capable engineering teams.
This is consistent with the wider promise of AI-enabled engineering. AI should augment engineering judgement, not create the illusion that judgement is no longer required. The organisations that gain most will be those that combine the speed of the tools with the architecture, validation, restraint, and stewardship of the profession.
Token economics deserves executive attention because consumption can accelerate far faster than traditional financial controls. The same acceleration applies to code changes, dependency acquisition, data access and command execution. The invoice tells you how quickly the agent was consuming a service. It does not tell you what the agent was doing, what it changed, or what authority it had while doing it.
The goal is not to slow every engineer down or to reserve all software creation for a protected elite class. It is to stop speed from being confused with control. Vibe coders and citizen developers should not be presumed to understand the full technical, cyber, data, and operational implications of a tool merely because they can use it to produce a working result. Professional engineers should not be presumed infallible either, but their discipline, supporting controls, and experience remain an essential part of responsible delivery.
If an organisation does not have that expertise internally, it should seek it. An independent consultancy or other qualified engineering and security specialists can help assess the tooling, its operating model, and the environment into which it will be introduced. That review is not bureaucracy for its own sake. It is how accountable leaders gain the information required to make a legitimate risk decision.
The organisations that benefit most from AI coding assistants will not be those that issue the most licences or burn the most tokens. They will be those that understand what authority they have placed behind the prompt, and those that engineer the environment around it accordingly.
Risks can be mitigated or accepted; but unknowns risks cannot be knowingly accepted. Before making a decision you are not qualified or experienced enough to make, consult the people who understand what is actually being placed at risk and follow their advice, or just role the dice. The decision is yours!

