Generative AI, large language models, and emerging agentic systems constitute the most disruptive transformation in the history of software engineering reshaping development processes, required competencies, professional roles, and the outcomes universities must deliver. A systematic review of 48 verified, influential peer-reviewed studies (2016–2026) synthesises the evidence along three trajectories: practice, education, and workforce. The corpus reveals a scientometric inflection annual LLM-for-SE output grew roughly five-fold after late 2022 and an evidence base internally contradictory on the size and even the sign of productivity effects. The central lesson: educating engineers for judgment, verification, and orchestration, rather than code production alone, is the defining challenge of the AI-native era.
Research activity accelerated sharply after publicly accessible generative AI arrived. Field-wide, the number of primary LLM-for-SE studies rose from 7 in 2020 to 13 in 2021, then jumped to 56 in 2022 and 273 in 2023 an inflection point that coincides with the late-2022 public release of ChatGPT and a parallel surge in computing-education research.
// primary LLM-for-SE studies per year ~5× rise from 2022 to 2023
// adoption climbed 76% → 84% of developers (2024→2025), yet self-reported trust fell the adoption–trust divergence.
The synthesis cuts across nine themes into three trajectories that move together. If entry-level coding becomes automatable, then what universities assess, what employers hire for, and what a "software engineer" means must all change at once.
Three tensions structure the evidence. They are not anomalies to be averaged away they are structural features of human–AI collaboration that any educational or organisational response must confront directly.
The synthesis organises into three mutually reinforcing pillars resting on durable computer-science foundations and bounded by an ethics-and-security envelope. Weak intent produces output that is harder to verify; weak verification makes collaboration unsafe; weak foundations leave the engineer unable to supervise the system at all.
The framework operationalises as nine competencies, each mapped to a dominant cognitive level. Where generation is cheap, the differentiating human contributions are those that judge, integrate, and direct so the model is weighted toward Evaluate and Create.
| ID | Competency | Cognitive level |
|---|---|---|
| C1 | Specification & intent engineering | Create / Evaluate |
| C2 | Critical evaluation of AI output | Evaluate / Analyze |
| C3 | AI-assisted debugging & verification | Apply / Analyze |
| C4 | Metacognition & self-regulation | Evaluate |
| C5 | Agent orchestration & tool use | Create / Apply |
| C6 | Foundational CS & systems thinking | Understand / Apply |
| C7 | Security, ethics & responsible use | Apply / Evaluate |
| C8 | Human–AI collaboration & communication | Apply / Create |
| C9 | Continuous learning & adaptability | Create |
A phased model protects deliberate practice early, then progressively shifts toward human–AI teaming and authentic, agentic projects. The unifying principle is assessment realignment: because code-writing tasks are now AI-solvable, assessment must privilege process, specification, evaluation, and the defence of work over artefact production alone.
THE
SCARCE
SKILL IS
JUDGMENT,
NOT
CODE.