Sooner or later, someone will ask IT a question like this: “What did the AI assistant show our sales team about the reorganization?” Or: “The automation sent a customer the wrong quote on Tuesday. What exactly did it do?” If you can’t answer in fifteen minutes, you have an observability gap, and AI makes it wider every month.
Most mid-size teams have decent logging for servers and sign-ins. Very few have thought about AI usage monitoring at an enterprise level: who is using which AI tools, what those tools are being asked, what data they touch, and what actions they take. That’s not a new discipline. It’s the same operational maturity you’d want for anything running in your environment, applied to a category of software that now makes decisions and takes actions.
Why this is an operations question
The assessment doesn’t ask about AI logging directly. It asks how automated and disciplined your IT operations are, because the teams that version-control their scripts and write runbooks are the same teams that can see, reproduce, and control what their AI does. The teams running on one person’s undocumented scripts can’t do either.
Here’s how the assessment asks the question, and the 0 to 4 ladder I score it against:
I3. How automated are your IT operations?
- Mostly manual; little scripting
- Individual scripts, not shared or version-controlled
- Shared scripts or an automation/RPA tool in limited use
- Automation is standard: version control, runbooks, an integration platform
- Infrastructure as code, CI/CD pipelines, and API-first integration are the norm
AI automations and agent configurations are operations code now. A prompt that drives an automation, the permissions an agent holds, the instructions a custom assistant follows: each of these belongs in the same discipline as your scripts. At level 1, they live in someone’s head. At level 3, they’re in version control, changes are reviewed, and there’s a runbook for when they misbehave.
What to log
You don’t need every detail on day one. You need enough to answer the questions that come up after something goes wrong. The minimum set:
- Who used which AI tool, and when. If your approved tools sign in through SSO, your identity platform already records this.
- What was asked and answered. Many enterprise assistants record interactions in an audit log. In Microsoft 365, for example, Copilot interactions appear in the Purview audit log. Check what your platforms capture, how long they keep it, and whether your license affects either.
- What data the AI touched. Which files or records were referenced in an answer. This is what you need when someone asks what the assistant showed a user.
- What actions it took. For agents and automations, every API call: what was created, changed, or sent, and under which identity. That’s why every integration needs its own identity, as covered in APIs before agents.
- What it was stopped from doing. DLP matches, blocked uploads, refused actions. These show where your controls are working and where people are pushing on them.
- What it cost. Token usage per application, which feeds straight into managing usage-based AI spend.
Where the logs should go
Send AI logs to wherever your security and operations teams already look: your SIEM or central log platform. A separate AI dashboard that nobody opens during an incident isn’t observability. If your AI tools can’t export logs to your central platform, note that as a limitation in your vendor review and keep those tools away from sensitive data until they can.
One caution: prompt logs contain whatever people typed, which can include sensitive information. Treat them as sensitive data. Restrict who can read them, set a retention period that matches your policy, and tell employees that approved AI use is logged. That notice belongs in your AI acceptable use policy, and it tends to improve behavior on its own.
The asset: five questions you should answer in 15 minutes
Use these as a test of your AI observability. Try each one this week. Any you can’t answer in fifteen minutes is your next piece of work.
- Who used our approved AI assistant last week, and how often? (Identity and usage reports.)
- What did a specific agent or automation change yesterday? (Action logs per integration identity.)
- Which documents did the assistant surface to a specific user this month? (Interaction audit logs.)
- Which AI application cost the most this month, and why? (Token usage per application.)
- Which unapproved AI sites did people use, and how much? (Web proxy, DNS, or browser logs.)
The fifth one is often the most revealing, and it’s the subject of auditing your own logs for shadow AI.
Runbooks for the three AI incidents you’ll actually have
Level 3 includes runbooks. For AI, write three short ones before you need them:
- The assistant surfaced something it shouldn’t have. Who confirms what was shown, who fixes the underlying permission, and whether it’s a reportable disclosure.
- An agent or automation took a wrong action. How to disable it, how to identify every action it took, how to reverse them, and who tells affected customers or staff.
- AI spend spiked. Who gets the alert, how to find the workload, and how to throttle or stop it.
Each runbook fits on a page. The act of writing them usually reveals the gaps in your logging, which is half the point.
Put AI configuration under version control
This is the part most teams miss. The system prompt for a custom assistant, the instructions for an agent, the configuration of an automation: when these change, behavior changes. If they’re edited in a web interface with no history, you can’t answer “what changed before the problem started?” Store them in version control alongside your other scripts, and review changes the same way. It’s the level 3 habit applied to AI, and it’s usually a small lift for whoever is already doing your automation work. If nobody is, that’s a skills gap covered in growing one AI builder on staff.
Mistakes I see at this stage
Assuming the vendor logs everything. Many do, some don’t, and retention varies. Check before you need it, not during an incident.
Logging everything forever. Prompt logs are sensitive. Keep them as long as your policy requires, and no longer.
Logs nobody reviews. A monthly look at the five questions above is enough to catch drift. No review means you’ll learn about problems from users.
Monitoring only the approved tools. The approved tools are the ones you already control. The risk is in the rest, which is why the fifth question matters.
What moving up one level looks like
From 0 or 1 to 2: get your existing scripts and any AI automations into a shared repository, and switch on the audit logging your approved AI tools already offer. From 2 to 3: route those logs to your central platform, write the three runbooks, and require review for changes to AI configurations. From 3 to 4: deploy automations and agent configurations through a pipeline, so every change is tested, recorded, and reversible. Each step is a quarter of part-time work for a mid-size team, and each one makes the next AI project cheaper to run safely.
Where does your team actually stand?
Operations automation is one of 24 questions in the AI Readiness assessment, which covers six dimensions: data, security, infrastructure, skills, use cases, and governance. The free version is 10 questions and gives you a score in a few minutes.
Get your free AI Readiness Score →
Want to see what the full assessment covers first? Flip through a complete 38-page sample report.
Related guides
- APIs Before Agents: Giving AI Access to Your Systems Safely
- Budgets and Alerts for Usage-Based AI Spend
- Publish Your AI Acceptable-Use Policy This Month (3-Tier Template)
- Shadow AI Statistics Are Scary. Your Own Logs Are Scarier
- You Need One AI Builder on Staff. Here’s How to Grow One
Frequently asked questions
What should we log about AI use?
Who used which AI tool and when, what was asked and answered where the platform records it, which data the AI touched, what actions agents and automations took, what controls blocked, and token usage per application. That set answers most questions that come up after something goes wrong.
Does Microsoft 365 record Copilot activity?
Yes. Copilot interactions appear in the Microsoft Purview audit log. Check how long your tenant retains those records and whether your license affects what is captured before you need them in an investigation.
Are AI prompt logs sensitive?
Yes. Prompt logs contain whatever people typed, which can include confidential information. Restrict who can read them, set a retention period that matches your policy, and tell employees in your acceptable use policy that approved AI use is logged.
Why should AI prompts and configurations be in version control?
Because behavior changes when instructions change. If a system prompt or an agent's configuration is edited in a web interface with no history, you can't answer what changed before a problem started. Version control and change review close that gap.




