Skip to content
    AppendicesChapter 132

    Integrating AI into Software Development at a Large Company

    State-of-the-art and production practices: engineering jobs to be done, AI in the SDLC, and measuring impact at scale.

    Original guideadvanced35 minvolatile · reviewed Aug 13, 2026
    Tech Lead
    Staff+
    Engineering Manager
    Chapter outline

    Brief

    The essential idea

    This approximately one-hour CTO Conf X talk argues that enterprise AI works when it is attached to concrete engineering jobs to be done rather than introduced as a universal assistant. Relevant jobs include API design, architecture drafts, boilerplate and onboarding, code review, tests, vulnerability checks, migrations, anomaly detection, on-call support, incident triage, and root-cause analysis.

    The talk places copilots and vibe coding beside MCP and agent protocols, while emphasizing enterprise integration and security limits. Platform investment and developer experience reduce the cognitive burden of adoption; without shared context, guardrails, and access to corporate tooling, promising demos do not become repeatable production practice.

    AI initiatives should follow an objective-metric-experiment model. Establish a baseline, run controlled rollout by team or domain, measure quality, speed, stability, cost, risk, and human time, and scale only effects that remain reproducible. Examples mentioned include Google API design, ByteDance review, Uber migration, Booking, Datadog observability and on-call support, and T-Bank's Nester copilot.

    Decision lens

    Key takeaways

    Start with a specific engineering job and constraint before selecting an AI capability.

    Build, quality, and operations contain different use cases and risk profiles.

    MCP and agent protocols still require enterprise security and integration design.

    Platform and DevEx investment make adoption repeatable across teams.

    A baseline and controlled rollout are necessary to distinguish real effect from novelty.

    Quality, speed, stability, cost, risk, and human time should be measured together.

    Scale only AI scenarios whose benefit is stable and reproducible in production.

    Workplace experiment

    Apply it at work

    1. 1

      Select one or two high-cost SDLC jobs and record a baseline for time, quality, and operational impact.

    2. 2

      Run a controlled rollout across defined teams or domains with an explicit comparison group where feasible.

    3. 3

      Evaluate model cost, security risk, human review time, and production stability alongside speed.

    4. 4

      Integrate the successful workflow into the developer platform with standard access and quality gates.

    5. 5

      Stop or redesign scenarios whose effect cannot be reproduced after the pilot.

    Choose one action, define the observable effect, and keep the first test small enough to reverse.

    Evidence

    Sources and further reading

    Primary source

    Additional sources

    Channel, aggregator, and commentary links confirm the work; they are not the primary source.

    Local knowledge map

    A small, typed neighborhood instead of the full-catalog graph.

    Previous chapterFrom Products to JTBD: How the Flow of Change Shapes Company StructureNext chapterFundamentals of Enterprise Architecture: Short Summary