The main principle is simple: it is possible to delegate the execution of an operation, but not the engineering judgment about which task to solve, which risk is acceptable, and whether the result may be used. All else being equal, the harder an error is to detect and the more costly its consequences, the more deeply a specialist should understand the subject matter and the more independent the verification should be.
This does not mean that everything must be written manually. On the contrary, programming and generative AI are useful precisely as means of automation. But automation changes how work is performed; it does not eliminate the need to assign an owner of the decision.
Four layers of competence serve different functions
Fundamental knowledge, subject-matter specialization, programming, and AI do not form four competing baskets of learning time. They answer different questions.
Fundamental knowledge provides the models through which a specialist explains a system’s behavior: causal relationships, constraints, invariants, ways of measuring, and characteristic failure modes. Knowing the fundamentals deeply does not mean remembering every formula and API. It means being able to reconstruct the reasoning and notice a result that contradicts how the system works.
Subject-matter specialization connects general models with a specific context. It includes the meaning of data, industry constraints, the real consequences of errors, and exceptions that are not visible in the code. An algorithm can be correct yet wrong for the subject matter if it optimizes the wrong metric or interprets the state of an object incorrectly.
Programming turns an intention into a precise, reproducible procedure. For some specialists, it is itself a deep professional field. For others, it is a tool for calculation, hypothesis testing, and automation. In both cases, it is necessary to understand the program sufficiently to assess its assumptions, observable behavior, and failure modes.
Generative AI accelerates the search for options, preparation of drafts, data transformation, and code generation. However, its output is not independent evidence of correctness. A model can produce a plausible result that does not correspond to the requirements, data, or architecture.
This leads to a useful distinction: fundamental and subject-matter competence determine what must be true; programming helps express and verify it in executable form; AI suggests options for how it could be done.
The depth of learning is determined not by the prestige of a topic but by the cost of misunderstanding it
It is rational to study more deeply where your own understanding is needed to formulate the task or detect an error. In practice, three levels can be distinguished.
| Required level | When it is needed | What the specialist must be able to do |
|---|---|---|
| Deep mastery | The error is hidden, the consequences are significant, the requirements are ambiguous, or the decision determines the architecture | Explain the model and constraints, derive correctness criteria, predict failure modes, and conduct a subject-matter review |
| Verification-level mastery | The task is standardized, but the result affects a working system | Read and critique the solution, check assumptions, run independent checks, and localize an error |
| Operational familiarity | The work is reversible, the harm is limited, and correctness is cheap to verify | Formulate a request correctly, compare the result with a reference, and safely roll back the change |
This allocation is not fixed once and for all. For example, syntax rarely requires continuous deep study if it can easily be checked by a compiler and tests. But a model of concurrent execution, data-consistency rules, or authorization semantics may require deep understanding: an error in these areas often passes local tests and appears only under a particular load or sequence of events.
A good way to test your own competence is to ask not “will I be able to get a working result?” but “will I be able to explain why it should work, under what conditions it will stop working, and what observations would reveal that?” If independent verification depends on asking the same AI again, independent understanding is probably still insufficient.
The boundary of delegation follows risk and verifiability
In risk management, AI risk is usually considered through the probability of an adverse event and the severity of its consequences. In practice, these quantities are often assessed qualitatively rather than calculated as exact numbers. NIST AI RMF recommends linking the use of AI to explicit risk management, measurement, and documented oversight, rather than treating the presence of a human as sufficient by itself (NIST AI RMF 1.0, AI RMF Playbook). Therefore, instead of pretending to have a precise numerical assessment, it is useful to begin with a qualitative checklist:
Before delegating, it is useful to assess seven properties of the task:
- Reversibility. Can the result be safely undone?
- Harm radius. How many users, data sets, and systems will an error affect?
- Observability. Will the error appear immediately or remain unnoticed for months?
- Cost of independent verification. Is there a test, reference, formal verification, or competent reviewer?
- Subject-matter uncertainty. Are the requirements unambiguous, and are the exceptions known?
- Data sensitivity. Is it permissible to transmit the context being used to the selected tool?
- Possibility of a safe rollback. Is there a tested recovery procedure, rather than only a theoretical rollback button?
Low harm, high reversibility, and inexpensive verification make it possible to delegate a substantial part of execution to AI. When observability is weak, consequences are irreversible, or requirements are ambiguous, AI is better used to search for alternatives and prepare a draft, while design and the final decision remain with a competent specialist.
Code complexity by itself decides little here. A large test-data generator may be safer than a short change to an authorization condition. The former is easy to isolate and verify, while one error in the latter can grant access to the wrong subject.
Generation speed is not the same as productivity or learning
Empirical results on the use of AI in development are heterogeneous. This is expected: studies differ in their tools, participants’ experience, task types, and selected metrics.
In a combined estimate from three field experiments with 4 867 developers, access to an AI assistant was associated with approximately a 26% increase in the number of completed tasks (estimate 26,08%; standard error 10,3 percentage points). Less experienced developers used the tool more often and saw a larger increase (The Effects of Generative AI on High-Skilled Work). This is a result under specific organizational conditions, not a universal estimate of AI’s effect on architectural quality, the number of defects, or any other team.
In another randomized study, 16 experienced developers performed 246 tasks in large open-source repositories familiar to them. With tools from early 2025, they worked 19 % longer on average, although they subjectively expected to be faster (METR study). The small participant group and specific scenario do not allow the result to be generalized to all development, but they show one important thing: the feeling of speed is unreliable.
Therefore, a team should measure the effect in its own workflow. Coding time is only part of the cost. Time spent formulating the task, reading the generated solution, making corrections, reviewing, integrating, and addressing subsequent defects should also be included.
A separate problem arises when AI is used not for production but for learning. In an experiment on learning a new Python library, the AI group received an average score of 50 % on a comprehension test versus 67 % for the hand-programming group, or 17 percentage points less. At the same time, the execution speedup did not reach statistical significance, and the result depended on how participants interacted with the assistant (How AI assistance impacts the formation of coding skills). This is a limited learning scenario, not evidence of harm from any AI. But it supports a reasonable precaution: when learning new material, do not delegate the very mental operation that must be learned.
If the goal is to learn how to design database queries, it is useful to ask AI to create test records or explain an error message. If AI instead immediately supplies a ready-made query and the learner only runs it, the learner develops the skill of obtaining an answer, but not necessarily a model of how the query is executed.
What can specifically be delegated to AI
In tasks with limited risk, AI is well suited to mechanical and draft work:
- creating scaffolds and repetitive code;
- converting formats and preparing test data without sensitive information;
- generating test variants from already formulated requirements;
- finding places that require attention during review;
- explaining an unfamiliar API followed by verification against the documentation;
- preparing implementation alternatives for engineering comparison.
Requirements definition, architectural boundaries, access models, irreversible data migrations, security, interpretation of ambiguous subject-matter rules, and release decisions should be delegated with caution. A ban on using AI is not mandatory here. Its role changes: from executor, it becomes a source of options evaluated by a specialist.
The more autonomous the tool, the more significant this distinction becomes. An assistant that suggests a code fragment has a smaller radius of action than an agent that runs commands, changes infrastructure, and publishes the result. Expanding authority requires stricter environmental constraints, prior approvals, and stopping points.
The result must be checked, not the generator’s confidence
The NIST SP 800-218A profile for generative AI does not divide source code into “human” and “AI-created”: it assumes that any source code must be checked for vulnerabilities and other problems before use (NIST SP 800-218A). The origin of the code may affect the nature of additional attention, but it does not replace ordinary engineering acceptance.
A minimum verification loop consists of several different sources of evidence:
- Record the requirements before generation. Otherwise, it is easy to mistake an attractive solution for a correct formulation of the task.
- Check the assumptions. Library versions, the data model, environmental constraints, permissions, and expected load.
- Read the change. Not only the final file, but also the diff, new dependencies, error handling, and removed checks.
- Run independent checks. Tests, static analysis, dependency scanning, and security checks proportionate to the risk.
- Conduct a subject-matter review. A technically correct implementation must correspond to the meaning of the data and business rules.
- Define post-deployment observation. Metrics, logs, an error signal, and a rollback procedure.
No single item proves absolute correctness. Tests check only specified cases; static analysis does not see every architectural error; code review depends on competence and available time. The strength of the loop comes from combining independent methods.
Asking the same model to “check its own answer” can be useful for finding obvious shortcomings, but it does not count as independent verification. The model may repeat the original incorrect assumption in different wording.
A formally designated human in the loop may change nothing
Human oversight is effective only when the designated person:
- understands the system’s purpose and limitations;
- receives enough information to verify it;
- has time, rather than mechanically approving hundreds of results;
- is able to recognize an error;
- has the authority to stop use of the result.
This understanding of oversight is also reflected in NIST recommendations on human-AI interaction (AI RMF, Appendix C). Requirements for human oversight of certain high-risk systems are formalized, for example, in Article 14 of the EU AI Act, but they cannot automatically be transferred to every system or jurisdiction (official text of Regulation (EU) 2024/1689).
Automation bias – the tendency to accept an automated recommendation without adequate verification – is particularly dangerous. A systematic review links this effect to trust in the system, user experience, workload, and verification complexity (Automation bias and verification complexity). Therefore, a confirmation button does not create meaningful control. Sometimes it merely transfers the signature to a human without giving that person a real opportunity to check the decision.
Responsibility should be assigned by decision
Delegating code generation or data analysis does not transfer responsibility to AI. In engineering work, it is useful to explicitly assign owners for at least the following decisions:
- the task’s goal and acceptable level of risk;
- requirements, constraints, and whether the use of data is permissible;
- architecture and system boundaries;
- implementation correctness;
- independent verification and acceptance criteria;
- release, operational monitoring, and incident response.
One specialist can combine several roles, especially in a small team. What matters is not the number of people, but the absence of ownerless decisions. The NIST AI RMF explicitly connects risk management with clear roles, authorities, and lines of communication (AI RMF Core). This is organizational responsibility; specific legal responsibility depends on the law, contract, and structure of the organization.
Example: AI writes a data migration
Suppose an engineer is preparing a migration that merges duplicate customer records.
AI can be asked to draft SQL, create synthetic test data, and suggest ways to check the number of rows. But the specialist must determine personally:
- what exactly counts as a duplicate;
- which records take priority in a conflict;
- which relationships and audit logs must be preserved;
- whether loss of individual fields is acceptable;
- how to detect an incorrect merge after the migration.
If the migration can first be run on a copy, its results compared with subject-matter invariants, and fully rolled back, the permissible scope of delegation increases. If the change is irreversible, affects sensitive data, and leaves no reliable indication of an incorrect merge, the primary focus should not be the quality of the prompt, but the design of the procedure, independent sample verification, and recovery.
In this example, deep study is required for data semantics and migration properties. It is sufficient to know the syntax of a particular SQL dialect at the level of confident reading and verification if documentation and a test environment are available. Mechanical generation can be delegated. The decision to run the migration remains with the designated owner.
A practical selection algorithm
For a new task, it is enough to follow a short sequence:
- Formulate the goal and the consequences of an incorrect result without AI assistance.
- Assess reversibility, harm radius, observability, data sensitivity, and the cost of verification.
- Determine which knowledge is needed to formulate the task and recognize a hidden error. That knowledge should be studied deeply.
- Identify mechanical operations whose results can be checked independently and inexpensively. They can be automated by a program or delegated to AI.
- Define acceptance criteria, verification, and rollback in advance – before receiving a result that looks convincingly correct.
- Name the person or team that will make the final decision and respond to the consequences.
The final boundary does not run between human and machine, nor between “simple” and “complex” tasks. It runs between decisions for which independent verification proportionate to the risk exists and decisions where the specialist is not yet able to distinguish a plausible answer from a correct one. It is in the latter category that the depth of one’s own knowledge remains indispensable.