AI can already produce a function, an SQL query, a set of tests, a deployment configuration, an API description, or a draft operating instruction in a matter of seconds. But the speed of producing text is not the same as the speed of building a working system.

Engineering work does not disappear along with the manual typing of artifacts. It shifts to what precedes generation and follows it: building a model of the task, formalizing requirements, choosing the solution mechanism, defining constraints, independent verification, integration, and managing consequences. In other words, AI can generate an answer to a formulated request, but someone has to establish which request corresponds to the real problem, by what criteria the answer is acceptable, and what will happen if it turns out to be wrong.

This distinction sets the practical boundary of automation. The production of a verifiable intermediate artifact can be delegated widely. AI can also help formulate the goal, admissible states, and acceptance criteria. But the approval of those criteria, the acceptance of residual risk, and the authorization to apply the result must belong to a specific accountable owner.

An Artifact Is Not Yet a Solution

An artifact is a discrete work product: source code, a query, a data schema, a test, a configuration file, a specification, an architecture diagram, a runbook, or a report. A solution includes not only such results but also the context of their use:

  • requirements and assumptions;
  • interfaces and dependencies;
  • resource and access-right constraints;
  • behavior under failure;
  • verification and observability methods;
  • update, rollback, and recovery procedures;
  • owners of decisions and risks.

A function may compile and pass tests yet violate an architectural invariant. A SQL query may execute yet answer the wrong business question. A runbook may be written clearly yet propose an operation that is inadmissible in the specific environment. In all three cases the artifact looks finished while the solution remains wrong.

In systems engineering this distinction is expressed through the concepts of verification and validation. Verification answers the question of whether the implemented product conforms to the established requirements. Validation checks whether the product is fit for its intended use in its intended environment and whether it meets stakeholder expectations. These are different tasks, and successful verification does not replace validation (NASA Systems Engineering Handbook).

AI can participate in both: generating tests, looking for contradictions, modeling scenarios, assembling reports. But the sufficiency of that evidence also has to be assessed. A test written from the same incomplete requirement as the implementation can confirm the internal consistency of two equally wrong interpretations.

What Exactly Is Automated

Generative tools are especially strong where input and output can be represented as a local transformation: a description into code, a schema into types, an interface into documentation, an example into a test stub. The more the result depends on implicit organizational or domain context, the less value there is in generation alone.

Artifact What is convenient to delegate to AI What remains subject to engineering review
Code Draft implementation, templates, API conversion, refactoring of a local fragment Semantics, architectural constraints, security, concurrency, resource limits, failure handling
SQL Building a query from a description, translating between dialects, explaining the plan Business meaning, schema, permissions, cardinalities, execution cost, empty and unusual data, risk of modifying or disclosing data
Tests Generating examples, fixtures, mock objects, boundary cases Completeness with respect to requirements and risks, independence of the source of expected results (test oracle), realism of the environment
Configurations CI/CD, container, policy, and infrastructure templates Secrets, permissions, version compatibility, scale of consequences, deployment and rollback strategy
Requirements Draft structure, language unification, searching for ambiguities Stakeholder intent, feasibility, consistency, verifiability, priorities and exceptions
Architectural descriptions Listing options, diagrams, recording the accepted decision Component boundaries, invariants, trade-offs, failure model, and long-term evolution
Runbooks and technical texts Draft, summary, format conversion, terminology alignment Factual accuracy, applicability to the environment, safe order of actions, currency

This table does not divide work into “creative human” and “mechanical machine”. AI is capable of proposing architectures, finding defects, and formulating requirements. The essential point is different: a proposal does not establish that the chosen option serves the system's goal.

For technical text this is as important as for code. Consider an instruction for restoring a service after a time-out: it suggests repeating the operation without checking whether that operation finished in the external system. If the retry is not idempotent, the operator may create a duplicate action. When checking such an instruction, what matters is the initial state, the source of reliable information about the result, admissible actions, success indicators, and the stop condition. Clarity of wording does not replace these checks.

Software development is not limited to writing code: before release you need to define requirements, design interactions, and integrate components; after release the system has to be operated, maintained, and eventually retired. ISO/IEC/IEEE 12207 sets a general framework of processes for the full life cycle of software systems from conception to retirement. Generating an artifact speeds up part of the work but does not establish its fitness for the system as a whole.

Problem Formulation Becomes Part of the Control Mechanism

The phrase “make a report on active customers” is understandable to a person only because that person fills in the context. What does “active” mean: at least one purchase, a valid contract, a login in the last 30 days, or no overdue debt? As of which date is the report built? How should returns, test accounts, and merged profiles be handled?

An LLM also fills in missing details, but a plausible completion does not necessarily match an organization's rules. Research on hallucinations in generated code shows that models can invent APIs, dependencies, and facts or else mislink the available context; this is a systematic class of failures, although it does not mean that every result is wrong (Exploring Hallucinations in LLM-Generated Code).

Therefore task formulation must convert stakeholder intent into requirements that can be implemented and verified. A useful requirement does not merely sound reasonable. It must be sufficiently unambiguous, technically correct, achievable, consistent, and tied to a verification method. These properties and requirements traceability are discussed in NASA Systems Engineering Handbook.

For a report on active customers, formalization may specify:

  • the exact activity condition and time zone;
  • data sources and deduplication rules;
  • the permitted update lag;
  • rules for handling deleted and blocked records;
  • permissions for viewing personal data;
  • control cases and expected totals;
  • the query execution time limit.

After that, AI receives a verifiable specification against which the result can be evaluated independently of the generator. The difference is fundamental: a prompt controls generation, while a specification makes it possible to judge the result independently of the generator.

Choosing the Solution Mechanism Cannot Be Reduced to Choosing a Model

Not every task formulated in natural language requires an LLM or another ML model. The engineer chooses between an ordinary program, a database query, a rule, an optimization algorithm, a finite-state machine, a simulation, a statistical model, and an agent with access to tools.

The choice is determined by the properties of the task:

  • whether the same answer needs to be reproducible;
  • whether exact algorithmic rules exist;
  • whether a probabilistic error is admissible;
  • how important latency and computation cost are;
  • whether data can be passed to an external service;
  • whether the decision needs to be explained or reproduced;
  • whether the result can be independently verified;
  • what happens if the model or an external tool is unavailable.

For example, calculating tax from a fixed set of rules is conceptually different from classifying free text. An LLM can help translate the rules into code or explain them. Execution of formalized rules should rely on a fixed version of the rules and independent checks; a generated implementation can be used when its correctness has been confirmed in light of the risk.

The same applies to agents. The ability to call a database, modify a file, and run a deployment increases usefulness but at the same time broadens the consequences of an error. Then the object of design becomes not just the model but the entire loop: permitted tools, permissions, limits, logging, confirmations, time-outs, retries, and emergency stop.

The Engineer Defines the Space of Admissible States

Architectural thinking starts not with choosing a fashionable component but with the boundaries of the system. Which states are admissible? Which transitions are forbidden? What must remain true under any action of a component?

Such properties are often expressed as invariants:

  • the sum of debits does not exceed the confirmed limit;
  • the operation is idempotent and does not create a duplicate payment;
  • a service without the required role does not gain access to data;
  • after a partial failure the record remains either in the old or in the new consistent state;
  • deleting a user propagates to the defined derived data;
  • when the model is unavailable the system enters a limited mode instead of treating a guess as a fact.

Mathematics and algorithms are needed here not for the sake of manually solving exercises. They provide a language for defining properties: pre- and postconditions, invariants, transition graphs, failure probabilities, resource constraints, algorithm complexity. Formalization does not necessarily mean a complete mathematical proof. For one component, types and contracts are enough; another needs property-based testing, model checking, or formal verification of individual critical properties.

AI can propose an invariant or an implementation of a check. But the admissibility of a state follows from the domain model and the risk of the system, not from the statistical plausibility of the text. If the engineer does not understand the algorithm, data model, or interaction protocol, they lose an independent basis for evaluating the result.

Example: Correct SQL with the Wrong Meaning

Suppose there are tables customers, orders, and payments. The task is to output the number of customers who “made a purchase last month”. AI offers a query in PostgreSQL syntax:

SELECT COUNT(DISTINCT customer_id)
FROM orders
WHERE created_at >= date_trunc('month', current_date) - interval '1 month'
  AND created_at <  date_trunc('month', current_date);

The query is syntactically correct and will probably execute. But it does not answer several essential questions:

  • is a created order counted as a purchase, or only a paid one;
  • are cancelled orders and returns excluded;
  • in which time zone are the month boundaries determined;
  • does customer_id refer to a person, an account, or an organization;
  • can one order have several payments;
  • is it permissible to read all rows of the table in the production database;
  • does created_at reflect the moment of ordering or the recording of the event after a delay.

Even execution accuracy, that is, agreement between the execution result and a reference on a specific dataset, does not fully solve the problem: two semantically different queries can coincidentally return the same result. Exact match compares a query with reference SQL according to the rules of a particular benchmark, not necessarily character by character. It can reject equivalent queries; implementation errors can, conversely, credit queries with different semantics as matching. Neither of these metrics by itself establishes whether the business event “purchase” has been defined correctly (analysis of Text-to-SQL metrics).

A correct solution will require defining the business event “purchase”, linking it to the schema, checking permissions, examining the execution plan, and testing the query on characteristic boundary data. SQL generation is only one step.

Verification Must Be Independent and Proportionate to the Risk

TEVV – testing, evaluation, verification and validation – combines testing, evaluation, verification, and validation. It is not a single final barrier but a set of ways to obtain evidence about system behavior before deployment and during operation. NIST AI RMF recommends explicitly defining roles, constraints, measurement and monitoring methods, documenting evaluation results, and managing risk throughout the life cycle (AI RMF 1.0). The framework model is voluntary and by itself does not establish a mandatory level of control for each system.

Different methods answer different questions:

  • testing shows behavior on selected inputs;
  • static analysis looks for certain classes of defects without executing the program;
  • inspection compares the artifact against requirements, rules, and architecture;
  • demonstration and end-to-end checks show the operation of a connected system;
  • simulation explores scenarios that are hard or dangerous to reproduce directly;
  • formal methods establish selected properties relative to a model and assumptions;
  • validation in the target environment checks whether the system solves the right task under real constraints.

No method proves the absence of all defects. A formal proof relates to a fixed model and set of properties. Tests cover only the scenarios examined. Static analysis is limited to a set of rules. User evaluation can uncover wrong meaning, but not a hidden data race.

Passing tests is also not equivalent to an acceptable change. In a METR study, four maintainers of three repositories reviewed AI-generated pull requests for 95 SWE-bench Verified tasks. By the authors' estimate, roughly half of the changes that passed the automatic grader would not have been approved by the maintainers even allowing for variability in their decisions: tests do not surface all problems with functionality, impact on other code, and implementation quality. The study covered only three of the twelve repositories in the benchmark; the review proceeded without the usual CI, and the agents were not allowed to refine the changes after comments. The conclusion applies to this experiment and does not measure the suitability of all AI changes in a normal iterative development process (Many SWE-bench-Passing PRs Would Not Be Merged into Main).

For security, AI-generated code should be included in the usual secure software development life cycle rather than treated as a special category exempt from accepted checks. NIST SSDF ties vulnerability risk reduction to practices of organizational preparation, protecting software, producing secure releases, and responding to vulnerabilities (NIST SP 800-218).

The Cost of Error Determines the Boundary of Delegation

The complexity of syntax is a poor criterion for autonomy. A one-line change to an access policy can be more dangerous than a thousand lines of an isolated prototype. It is more useful to assess four properties:

  1. Verifiability. Can the correctness of the result be determined independently?
  2. Reversibility. Can the action be undone quickly and reliably?
  3. Observability. Will the error be detected before substantial harm?
  4. Scale of consequences. How widely will an error, leak, or outage spread?

From these follows a working delegation matrix. This is an engineering recommendation, not a universal standard.

Conditions Admissible mode Necessary evidence and constraints
The result is local, reversible, and easy to verify Automatic generation; automatic application only after mandatory automated checks that block application on failure Tests, static analysis, change log, limited permissions, and a verified rollback
Significant context is required, but the error is detectable and rollback is reliable Generation with automated checks; mandatory review of changes affecting interfaces, data, or permissions, spot review of isolated changes Traceability to requirements, integration tests, compatibility checks, rollback plan
The error affects data, security, availability, or many users AI proposes; an authorized person or an independent loop approves Risk analysis, separation of duties, testing in an isolated environment, monitoring, staged rollout
Consequences are hard to detect, localize, or compensate Autonomous application is usually unacceptable without a specially designed safeguard loop Formalized constraints, independent verification, fail-safe behavior, stop authority, audit

An exception to a requirement or constraint must have a justification, an owner, a review date, and a record of the residual risk. It must not be turned into an unnoticed permanent rule just because an automated process is able to continue running.

Human oversight does not necessarily mean line-by-line review of every file. It can be based on contracts, automated checks, spot inspection, and reinforced review of critical sections. What matters is that oversight is independent of the same incomplete context that produced the result, and that the authorities are defined in advance.

Integration and Operation Restore the System Context

Local generation usually cannot see the whole system. Between a correct fragment and a working release remain:

  • schema and API compatibility;
  • migration of existing data;
  • the order of component deployment;
  • secrets and permissions management;
  • behavior under partial unavailability of dependencies;
  • retries, time-outs, and idempotency;
  • metrics, logs, tracing, and alerts;
  • rollback and recovery after an incident;
  • documentation updates and operator training;
  • further maintenance and technical debt.

Technical debt here means the future cost of simplifications and inconsistencies: duplicated logic, blurred interfaces, fragile tests, hidden dependencies. AI can both reduce it through consistent refactoring and accumulate it through many locally reasonable but architecturally incompatible changes.

Operation adds information that does not exist at the generation stage: real input distributions, load patterns, rare failures, user actions, and interaction with external systems. That is why evaluation does not end at release. Significant systems need monitoring of assumptions, incident analysis, and re-verification when the model, agent, retrieval context, tools, or architecture change.

The provenance of an artifact also becomes part of engineering traceability. The version of the model and tools, the provided context, human changes, check results, and accepted exceptions are useful to keep as process work products. Such a record does not make a decision correct, but it makes it possible to reproduce the check, investigate a defect, and understand who accepted the residual risk.

Generation Speed Is Not the Same as the Productivity of the Development System

Studies of AI assistants give different results because they measure different tasks, groups of developers, and working conditions. In a randomized controlled experiment, 95 professional developers were split between a group with GitHub Copilot and a control group and asked to implement an HTTP server in JavaScript. In these specific conditions the average task completion time in the Copilot group was 55.8% lower; the 95% confidence interval for the time reduction reported by the authors ranged from 21% to 89% (The Impact of AI on Developer Productivity).

In another randomized study, METR included 16 experienced developers performing 246 tasks in mature open source repositories familiar to them. With the AI tools available in early 2025, estimated task completion time was 19% higher, with a 95% confidence interval from 2% to 39% increase in time (METR study; subsequent clarification by the authors).

These point estimates cannot be directly compared as universal effects of a single intervention. The studies differed in sample, task, codebase, tools, and measurement procedure, and the wide confidence intervals reflect statistical uncertainty within each experiment. Together the results show only a more general conclusion: the effect depends on the task, the maturity of the codebase, the experience of participants, the quality of context, and the cost of verification.

Therefore evaluating automation only by the number of generated lines, closed tickets, or time to the first patch is insufficient. For a system, the following also matter:

  • functional correctness;
  • the cost of review and rework;
  • defects and incidents after release;
  • security and performance;
  • integration failures;
  • the maintainability of the change;
  • recovery time;
  • the accumulation of technical debt.

A tool can speed up writing the first version while simultaneously increasing the volume of verification. In another context it can reduce both quantities. The conclusion should be drawn from the full workflow, not from the speed of typing an artifact.

How the Engineering Function Changes

When artifact production becomes cheaper, the value of an engineer is determined less and less by the speed of producing syntactically correct text. What matters more is preserving the causal chain from need to observable system behavior. To do this you need to:

  • clarify the goal and conflicting stakeholder expectations, then convert them into verifiable requirements;
  • choose the appropriate algorithm, model, and architecture, and define interfaces, invariants, and forbidden states;
  • evaluate the consequences of error and design proportionate checks, including verification, validation, and observation of the system in operation;
  • integrate components and manage changes, exceptions, and technical trade-offs;
  • distribute authority for release, stop, and rollback, and document the accepted residual risk.

Domain knowledge, mathematics, algorithms, and architecture do not become less necessary in the process. They turn from means of manual production into means of independent control. To spot a fictional API, you need to know the ecosystem. To reject a slow SQL query, you need to understand the schema and cardinalities. To check a model, you need statistics and knowledge of the data. To detect an architectural violation, you need to see the system beyond the local patch.

AI can be the author of a significant part of the artifacts and a participant in verification. But responsibility requires explicit roles: who approves requirements, who authorizes an exception, who accepts a release, who can stop the system, and who investigates an incident. NIST AI RMF treats such roles, risk measurement, and risk management as organizational functions rather than properties of a single tool.

The final boundary does not lie along the line “the human writes, the machine suggests”. It lies between generating a possible result and a justified decision about its applicability. As long as a system must serve a defined purpose, remain within admissible states, and withstand the consequences of errors, the engineer is responsible not for the volume of text written but for the coherence of the problem model, the implementation, the evidence, and the operation.