An AI-native engineering team uses coding agents such as Claude Code and Codex on its own repository, from a clear brief to tested changes, with review, evaluation and repository standards doing the work that a mandate cannot. The route TAI Labs runs is a 4 to 12 week programme built on the team's real codebase: a repository-aware agent setup, custom MCP connections to internal systems, evaluation sets for the AI features the team ships, and engineering standards packaged as skills the team owns. This guide is that plan. It is not a plan for measuring adoption per developer, and it says why.
What changes when an engineering team is AI-native
What engineers notice first is that the agent knows the repository: the build commands, the conventions, the boundaries, the tests that matter. With that context in place a coding agent plans a change, builds it, runs the relevant tests and hands a reviewer a focused patch with its assumptions listed. Without it, the agent produces plausible code that fails review, and the team concludes the tools do not work. Most of the programme is putting that context in place and teaching the review habits that go with it.
- A brief becomes a plan, a focused patch and the relevant tests before a human reads it.
- The agent's context (CLAUDE.md or the equivalent, skills, MCP connections) is versioned in the repository and maintained like code.
- AI features ship with an evaluation set, so a regression is a number rather than an argument.
- Code review uses AI to find failure paths and reproduce issues, and a person decides what ships.
Start with the workflows, not the tools
Engineering teams start with the change workflow below, on a bounded task in their own repository. The other ten, including the MCP server, the retrieval-backed assistant and the evaluation workflow, are on the AI training for engineering teams page.
Example workflowTurn a product brief into a working prototype.
01 · Start with your tools
Your brief, codebase, sample data and team standards.
02 · Apply your process
Plan the change, build with a coding agent and run relevant tests before review.
Build a reusable workflow or skill03 · Review the result
Working code, checks and a clear list of what to improve next.Your team checks and approves- Set up a repository-aware coding agent: build commands, boundaries and examples documented.
- Build a custom MCP connection to an approved internal API with scoped permissions and tested errors.
- Develop a retrieval-backed assistant that exposes sources and is evaluated where the answer is absent.
- Build a multi-step agent workflow with approval points and tested failure paths.
- Generate test cases that detect real problems, not just pass.
- Review a change with AI assistance and validate findings before reporting them.
- Modernise a bounded legacy component with rollback notes.
- Evaluate an AI feature against representative inputs and unacceptable failures.
- Turn engineering standards into versioned skills the team improves together.
- Prepare an incident investigation that separates hypotheses from facts.
The tools an engineering team actually uses
- Claude CodeAgentic development in the terminal on your repository
- CodexCoding agent, reusable skills and background tasks
- CursorAgent-assisted editing inside the IDE
- GitHub CopilotCompletions and review inside GitHub
- GitHubPull requests, checks and the review loop
- Custom MCPsScoped connections to internal APIs, tickets and docs
- LangGraphOrchestrate multi-step agent workflows and evals
- Hermes AgentAgents with reusable skills for repeated engineering tasks
- LinearPlans into tracked work
- JiraThe same, for teams on Jira
- GeminiLong-context reading of logs and documents
| Job | Claude Code | Codex | Cursor |
|---|---|---|---|
| Multi-file change from a brief, with tests | Strong: plans, edits, runs tests | Strong, including background runs | Good, with the editor in view |
| Repository context that persists | CLAUDE.md, skills, MCP | AGENTS.md, skills | Rules files |
| Connecting internal systems | MCP servers | MCP servers | MCP servers |
| Review of a pull request | Good | Good | Good, inline |
A plan to make the team AI-native, in 4 to 12 weeks
Engineering programmes run shorter than the others when the team already ships daily, because the labs happen inside the normal delivery cycle. WireApps ran as a live multi-time-zone cohort for more than fifty engineers; a single team can finish in four weeks.
Interview the engineers about where time goes
Week 1. Review rework, flaky tests, the incident that took a week to understand, the internal API nobody wants to touch. The bottlenecks decide the labs.
Set a baseline the engineers agree with
Time to a reviewed, tested change on a defined class of task; review rework and regressions; performance on an evaluation set for one AI feature. No per-developer adoption metrics.
Agree access, boundaries and review rules
Which repositories, which data, what an agent may and may not run, and the standard a patch must meet before a reviewer sees it. Security and the tech leads sign this off.
Lab one: the repository-aware agent
Weeks 2 to 3. The team documents build commands, conventions and boundaries for the agent, then every engineer takes a bounded brief to a tested patch with an instructor beside them.
Lab two: connections and retrieval
Weeks 4 to 5. A custom MCP server over an approved internal API, with scoped permissions and tested errors, and a retrieval-backed assistant with an evaluation set that includes the questions it should refuse.
Labs three and four: agents, evals and review
Weeks 6 to 9. A multi-step agent workflow with approval points, an evaluation of a shipped AI feature, AI-assisted review that validates its own findings, and the engineering standards packaged as skills.
Office hours
Throughout. Two open sessions for the agent that keeps ignoring the convention and the eval that passes for the wrong reason.
Demo day and handover
Weeks 10 to 12. Engineers show working changes, the MCP server and the eval reports. The agent context, skills and evals live in the repository; the programme itself is handed over on your learning platform with an executive readiness report.
The knowledge and training system the team keeps
For engineers the system lives in the repository, which is why it survives. The agent context, the skills and the evaluation sets are reviewed in pull requests like everything else, and the next engineer inherits them on their first clone.
- Agent context files per repository: commands, conventions, boundaries, examples
- Versioned skills for the engineering standards, with helper scripts
- MCP servers with documented access boundaries and tests
- Evaluation sets for each AI feature, run in CI
- A review checklist for AI-assisted changes
- The programme on your learning platform for the next intake, and the executive readiness report
Do it with us once
We build this with your engineering team, once.
Discovery, labs on your repository, the agent context, MCP connections and evals, and the handover. Bring one bottleneck to a 15-minute call and we will scope it. If you would rather look first, run the free [AI workflow audit](/ai-workflow-audit) or watch the [sample Cursor lesson](/custom-training#training-team).
Why we do not measure adoption per developer
Dashboards of AI usage per engineer, and mandates to use the tools, produce resentment and gaming, not better software. Engineers describe them as surveillance, and they are right that an adoption number says nothing about whether the code got better. The measures in this programme are about the work: time to a reviewed change, rework, regressions, evaluation scores. If the tools help, those move. If they do not, the team should know that too.
- No per-developer usage metrics are collected or reported.
- Agents run only inside the boundaries the team documented, and never with production credentials.
- A patch is reviewed by a person who can reject it, and the review checklist is part of the standard.
- AI features ship with evaluations that include the cases where the right answer is no answer.
- Retrieval, transcripts and documents are treated as untrusted input in every agent prompt.
How to measure whether it worked
- Time to a reviewed, tested change on a defined class of task, before and after
- Review rework and regressions per change
- Performance on the evaluation set for one AI feature, release over release
- Incidents where the investigation brief was ready before the on-call engineer opened the logs
Mistakes that stall engineering teams
- Buying licences without documenting the repository for the agent. The tools then look worse than they are.
- Mandating usage. It converts sceptics into saboteurs.
- Skipping evals. An AI feature without an evaluation set is an opinion in production.
- Giving agents broad credentials to save time. Scope permissions first, widen later.
- Letting the standards live in one senior engineer's head instead of in a skill the team reviews.
“Well structured, and covered a genuinely broad range of knowledge. Exactly what our engineers needed.”

Frequently asked questions
Can the programme work with our existing codebase?
Yes, subject to your access and data policies. We can also start with a representative sample repository, and we agree the stack, review requirements and scope before the sessions.
Does this cover more than prompting a coding assistant?
Yes. Depending on scope: repository context, custom MCP servers, agent orchestration, retrieval, evaluations, reusable skills and the review practices needed to maintain the work.
Claude Code or Codex?
Both are taught. Most teams standardise on one terminal agent and keep an IDE agent such as Cursor or GitHub Copilot alongside it. The choice depends on your stack and your model access, and we make it with you during scoping.
Will you report which engineers use AI the most?
No. We do not collect per-developer usage metrics. The readiness report scores the assessed work the team produced against the baseline.
How long does it take?
Four weeks for a single team with a clear bottleneck; up to twelve for a larger cohort covering connections, agents and evals. Dates and time commitment are agreed in the proposal.
What does it cost?
Indicative ranges for a 20-person team are on the custom AI training page, and the cost guide explains what moves the number. A written proposal follows the scoping call.
Give the agent the repository context it needs, scope its permissions, teach the team to review and evaluate what it produces, package the standards as skills in the repository, and measure the work rather than the people. That is what TAI Labs does with engineering teams in four to twelve weeks.