Ship production LLM systems responsibly — retrieval, evaluation, guardrails and agents. A 10-week live cohort built on the safety practice enterprises now require.
Anyone can call an API. Shipping something an enterprise will trust is a different discipline.
Shipping a generative-AI system that’s reliable, evaluated and safe enough for an enterprise to trust is a different discipline from calling an API — and it’s the one this program teaches.
Across ten live weeks you’ll build production LLM systems end to end: retrieval, evaluation, guardrails and agents, with the safety and governance practice enterprises now require before anything reaches their users.
Small cohort, senior-led, grounded in real deployments.
They score the generation and never score the retrieval underneath it, then tune prompts to fix a search problem.
Injection and exfiltration are not edge cases. You red-team your own build and report the attacks that worked.
An enterprise review asks for documentation, accountability and evidence. We produce those, not a paragraph about values.
Movement I · Weeks 01–03RetrieveMap the whole system, then build retrieval you can measure on its own.
Movement II · Weeks 04–06ProveTyped output, tool use, and an eval suite that gates every change.
Movement III · Weeks 07–10DefendGuardrails, red-teaming, contained autonomy and the governance record.
The systems viewA generative AI system is retrieval, prompting, tools, evaluation and guardrails together. We start by mapping the whole thing before touching any part of it.
A system map of your build, and a cost ceiling set before you write code.
Retrieval that worksChunking, embeddings, hybrid search and re-ranking — and how to measure retrieval quality separately from generation quality, which is where most teams go wrong.
A retrieval score you can quote, measured independently of the model on top.
Structured output and tool useGetting reliable, typed output from a probabilistic system, and letting it call tools without losing control of what it does.
Typed output and tool calls that fail loudly rather than quietly.
EvaluationBuilding an eval suite that reflects your actual task: golden sets, model-graded evaluation and its limits, regression testing on every change.
A golden set and a regression gate that runs on every change.
Guardrails and adversarial behaviourInjection, jailbreaks, data exfiltration and the failure modes specific to systems that read untrusted text. Red-teaming your own build.
A red-team report on your own system, including the attacks that worked.
Agents, carefullyMulti-step autonomy, where it genuinely helps, and the containment that has to exist before it ships.
Containment you designed before the autonomy, not after an incident.
Governance and the AI6 practicesDocumentation, accountability and the record an enterprise will ask you for. What responsible deployment looks like in artefacts, not adjectives.
The governance artefacts an enterprise security review actually asks for.
Your systemBuilt, evaluated, red-teamed and documented to a standard an enterprise review would accept.
A production-grade LLM system you would put in front of a security review.
You build a generative AI system against a real corpus and take it to a standard you would be willing to put in front of an enterprise security review: measured retrieval, an evaluation suite, guardrails, and a written red-team report on your own system.
The presentation includes the attacks that worked. Every system has some; the assessment is on how you found them and what you did next.
Advanced. You should have shipped software professionally and be comfortable with Python, APIs and version control. AI Engineering & Production or equivalent production experience is the expected starting point.
Ten weeks live, roughly ten hours a week outside sessions. You need a cloud account and API access; we cover cost control in week one because this is the program where an unattended loop gets expensive.

Frameworks are great. But understanding what they abstract — the loop, the parsing, the failure modes — is what lets you actually debug and design them under pressure.
Amir founded, funds and directs AI Tech Institute, and still teaches. He holds a PhD and has spent more than twenty years building and leading AI, machine learning and data teams in Australian industry — the kind of work the programs here are drawn from rather than adapted to.
He currently leads asset performance and analytics at Synergy, Western Australia's state-owned electricity generator and retailer. Before that he led GenAI and MLOps program delivery at EY, working on platform modernisation for tier-one Australian banks; served as Chief AI Engineer at Hancock Prospecting, where he built enterprise GenAI platforms for executive decision support, led a Snowflake migration, stood up on-premises NVIDIA H100 infrastructure and Kubernetes clusters for production AI, and built a team of twelve engineers; and spent nearly five years as Tech Lead for ML and Data Science at Woodside Energy.
He has been teaching for twenty-one years. He is an Adjunct Associate Professor at the University of Western Australia, teaching quantitative analysis and decision-making at master's level, and holds Databricks University Alliance Faculty status and the AWS Certified AI Practitioner certification. He has published peer-reviewed research in Exploration Geophysics on machine learning for reservoir characterisation, and writes and records openly about the parts of AI engineering most courses skip — agent loops, termination logic and failure modes.
Cohorts are deliberately small so every piece of work gets reviewed properly and every presentation gets heard in the room. If an intake is full we’ll say so and hold your place for the next one — we won’t quietly oversell a cohort.
See all upcoming dates →The prerequisites for this program are in “Who it’s for” above. Our foundation-level programs and workshops assume none. If you’re unsure, ask on a qualification call — we would rather redirect you than take your money for the wrong cohort.
The commitment for this program is shown in the hero and in “Who it’s for”. Cohort programs run live sessions plus practice between them; workshops are a single block with no homework.
The work you shipped — defended in front of the cohort — and an honest recommendation on your next step. Sometimes that is another cohort with us, sometimes it isn’t; we will tell you which.
No. AI Tech Institute is not a Registered Training Organisation and this is not a nationally accredited qualification. It is professional education, assessed on the work you ship, with a certificate of completion from us. We won’t claim accreditation we don’t have.
Yes — we invoice organisations directly and can supply a scope and outcomes summary for your L&D or capability budget. If three or more people from one team want in, talk to us about a private cohort instead.
Full refund any time before the program starts. Once it has started we can’t refund the place, but we will carry you into a later intake where the circumstances warrant it. Ask us — we’re reasonable.
Thirty minutes with us and you’ll know whether this is your right next step — or which of our programs is. No pitch, no payment, and we will say so if the answer is “not this one”.