Service · Audits & reviews

AI & software system audits

An AI system audit is a structured technical review that answers four questions: where does your system break, why, how often, and what will it take to make it dependable? NorthSight Technologies performs audits of LLM pipelines, AI features and the software platforms around them — producing a prioritized, evidence-based roadmap instead of a vague health check. Start with a free introductory call.

TL;DR
  • An audit covers architecture, failure modes, security & data handling, evaluation coverage, and cost.
  • You get a written findings report, a failure taxonomy, and a remediation roadmap ordered by impact.
  • Common triggers: AI that works in demos but fails in production, runaway LLM costs, inherited systems nobody fully understands, upcoming compliance reviews.
  • Start with a free 20-minute call to scope the audit around the failure that hurts most.
By Kamran Khalil, Technical Architect & Co-founder · NorthSight Technologies, Ontario, Canada · Published July 7, 2026

What an audit actually covers

"Audit" gets used loosely. Ours is a hands-on engineering review with five concrete lenses:

  • Architecture. How the system is put together: data flow, model choices, integration points, single points of failure, and whether the design matches the problem — or fights it.
  • Reliability & failure modes. We build a failure taxonomy from real traces: hallucinations, bad tool calls, loops, context blowups, silent data corruption. What breaks, where, and how often — measured, not guessed.
  • Security & data handling. Where data actually goes (including prompts, logs and caches), access control, residency boundaries, and exposure to prompt-injection and data-leakage paths. For Canadian public-sector work this includes Protected-B and data-residency considerations.
  • Evaluation coverage. Whether quality is measured at all — and if so, whether the evals reflect real usage. Most AI systems we review have no regression safety net; that is usually finding #1.
  • Cost. Token spend, model-tier choices, caching opportunities and architectural waste. LLM bills often carry 30–70% avoidable cost from oversized models and redundant calls.

What you receive

  • A written findings report — specific, evidence-backed, readable by both engineers and leadership.
  • A failure taxonomy of the observed and probable failure modes, ranked by frequency and blast radius.
  • A prioritized remediation roadmap — ordered by impact, each item scoped so you can act on it with us or with your own team.
  • A working definition of "good enough to ship" for your context: the measurable bar the system must clear before you can stand behind it.

The audit is deliberately vendor-neutral: the roadmap is yours, and it is written so any competent team can execute it. Clients often engage us for the fixes, but the report does not depend on that.

Signals you need one

  • An AI feature that impressed everyone in the demo and now misbehaves with real users.
  • An LLM bill growing faster than usage, with nobody able to say exactly why.
  • A system inherited from a departed team, an agency, or an acquisition — capability unknown.
  • A security, procurement or compliance review on the calendar.
  • You are about to scale usage and don't know if the system will hold.

How it runs

This is step one of our Reliability-First Method. We start where it hurts most: you show us the failure that costs you the most, we trace it to root cause, and we review outward from there. Start with a free 20-minute call — a low-risk way for both sides to see whether a deeper engagement makes sense. Audits are then scoped based on system size and access requirements.

Frequently asked questions

What is an AI system audit?

A structured technical review of an AI feature or pipeline — architecture, failure modes, security and data handling, evaluation coverage and cost — producing a prioritized, evidence-based list of what to fix first.

How is this different from a code review?

A code review looks at code quality line by line. A system audit looks at behaviour: what the system actually does in production, where it fails, what that costs, and which fixes matter. Code quality is one input among several — traces, evals, architecture and data flow matter more for AI systems.

How much does it cost?

Audits are scoped as paid engagements, based on system size and access requirements. Your first 20-minute call is free — we use it to understand the failure that hurts most and outline how we'd approach the audit. Email hello@northsight.ca to start.

Do you audit non-AI software systems too?

Yes. The same lenses — architecture, failure modes, security posture, cost — apply to conventional web platforms, mobile apps and backend systems. Our principals have 28+ years of combined engineering across government, enterprise and consumer software.

What do you need from us to start?

For an initial conversation: the failure that hurts most, plus whatever you can share — error examples, traces, screenshots. For a full audit: read access to the relevant repositories, logs and dashboards, under NDA as needed.

Start with the failure that hurts most.

Book a free call — we'll trace one real failure to root cause and tell you honestly how we'd approach the audit. Free 20-minute call, no pitch.