HR automation is reliable not when a model performs an isolated task well, but when the whole process withstands its errors. The key question, therefore, is not “can AI rank resumes or compose an employee evaluation”, but “what action follows the output, who will check it, what data and criteria it rests on, whether it can be challenged, and who is responsible for the consequences”.
The practical boundary runs along the impact on a person:
- Administrative and verifiable operations are usually suitable for automation: field extraction, status cross-checks, request routing, search across approved policy, reminders and draft preparation.
- Decision support is permissible only with a valid criterion, controlled data, measurable quality and a person with real authority to change the result.
- Decisions about access to work, income, promotion, dismissal, shifts or performance evaluation must not be turned into an autonomous pipeline merely because it is technically possible.
- Some applications may be prohibited by law regardless of the claimed effectiveness.
This boundary does not coincide with the product name. A chatbot may be a low-risk reference tool while it quotes the approved leave policy. The same interface becomes part of a significant decision if it interprets medical information and recommends refusing an employee special working conditions.
What exactly is automated
Under the word “automation” different objects are often hidden:
- Workflow – the system routes a request to the right team, checks completeness or triggers a reminder.
- Extraction and classification – the model finds fields in a document or determines the topic of an enquiry.
- Generation – the model composes a letter, a job description, a reply to an employee or a summary of an interview.
- Forecast or ranking – the system assigns a probability, a score or a place in a list.
- Decision and action – the system rejects a candidate, assigns a shift, restricts access or launches a disciplinary procedure.
The last two objects differ fundamentally from the first three. An error in an extracted date can be detected by comparison with the document. An erroneous “suitability” score has no equally obvious benchmark: it depends on the definition of suitability, historical decisions, the chosen horizon and which outcomes are observable at all.
Automated decision-making means that the outcome is formed by system means without substantive human involvement. A formal confirmation button does not necessarily change this qualification if the reviewer usually accepts the recommendation without independent analysis.
Human-in-the-loop – a process in which a person is not merely present, but understands the purpose and limitations of the system, sees the necessary grounds, and has the time and authority to stop or change the action.
Applicability map: from useful automation to prohibition
| Application class | Examples | Reasonable position |
|---|---|---|
| Verifiable administrative operations | Extracting record details, checking package completeness, routing enquiries, reminders | Automate with sampling checks, logging and secure data handling |
| Search and draft preparation | Search across approved HR policy, a draft letter or an onboarding plan | Automate with source attribution and review before sending |
| Aggregated operational analytics | Enquiry volumes, processing times, unclosed tasks | Automate if covert individual evaluation and re-identification are excluded |
| Screening and ranking of candidates | Filtering resumes, shortlist, matching with a vacancy | Only after validating the criterion, checking for discriminatory effect and organising review |
| Employee evaluation and management | Performance score, promotion recommendations, allocation of shifts and tasks | High-risk loop: separate justification, strict control and limits on autonomy are required |
| Sanctions and termination of the relationship | Automatic rejection, dismissal, reduction of compensation, disciplinary measure | Do not hand over to the system as an autonomous action |
| Inferring emotions in the workplace from biometric data | Determining “engagement” from face, voice or other biometric signals | Prohibited in the EU, except narrowly framed medical and safety exceptions |
The classification refers to the actual use, not to the marketing description. For example, a “hint to the recruiter” in fact determines the outcome if candidates below a threshold are never reviewed or if the recruiter is shown only the top of the list.
The legal boundary depends on jurisdiction and purpose
In the EU, Regulation (EU) 2024/1689 classifies as high-risk a number of systems for hiring and selection, decisions on terms of the employment relationship, promotion and termination of the relationship, allocation of tasks on the basis of behaviour or personal characteristics, as well as employee monitoring and evaluation. The reason is the potential substantial impact on career, livelihood and people’s rights. The exact status is determined by the purpose and manner of use of the specific system, not by whether its function is called administrative (text of Regulation (EU) 2024/1689, explanation of recital 57).
For high-risk systems the regulation provides for a connected set of measures: risk management, data requirements, technical documentation, logging and traceability, transparency, human oversight, accuracy, resilience and cybersecurity. The deployer, that is the organisation applying the system, retains obligations to use it in accordance with the instructions, control input data, monitor it and appoint competent persons for oversight (European Commission overview, Article 26). A contract with the vendor does not transfer the whole of the employer’s responsibility to it.
The AI Act also prohibits systems intended to infer the emotions of a specific person in the workplace on the basis of biometric data, except for medical or safety purposes (Article 5). This prohibition should not be extended to any satisfaction analytics: an anonymous survey or an aggregated organisational indicator is a different object of processing. But renaming a biometric inference an “engagement score” does not change its essence.
Personal data protection applies separately. Where applicable, the GDPR restricts decisions based solely on automated processing that create a legal or similarly significant effect. The exceptions are narrow and accompanied by safeguards, including the possibility of human intervention, expressing one’s position and challenging the decision (EDPB on data subject rights, ICO guidance). The presence of human review does not by itself make the processing lawful: the review must be substantive, and the purpose, legal basis and scope of data must be justified.
As of 20 September 2026, the official European Commission page indicates 2 December 2027 as the date when the AI Act rules begin to apply to high-risk systems in the area of employment. This date differs from the schedule that can be derived from the original version of the transitional provisions of Regulation (EU) 2024/1689 without taking subsequent amendments into account. Therefore, for a practical assessment it is necessary to check the currently applicable consolidated version of the regulation and the specific transitional provisions, rather than relying only on the original text of Article 113. The prohibitions of Article 5, points (a)-(h), including the prohibition on inferring emotions in the workplace from biometric data, apply from 2 February 2025. The classification given is an example of EU law, not a universal rule for all countries.
In the US there is a different legal framework, but the practical conclusion is similar: an algorithmic tool that makes or informs a hiring, promotion or dismissal decision may be treated as a selection procedure. Overall accuracy does not answer the question of adverse impact on protected groups; what matters is the relationship of the criterion to the job, business necessity and the availability of a less discriminatory alternative that is equally effective. This approach is reflected in EEOC materials on employment tests and selection procedures; the agency’s clarifications are not themselves an independent rule of law. The legal assessment depends on the specific application and facts. These requirements cannot be mechanically mixed with European ones, but they show a common defect in the approach “the vendor reported high accuracy, therefore the system can be used”.
Why a technically working model can make the HR process worse
Unreliable ground truth
Ground truth – the observed outcome used as a benchmark for training or validating the model. In HR it often reflects not the objective quality of an employee, but the organisation’s past decisions.
If a model learns from historical manager ratings, it reproduces the criteria of those ratings, including inconsistency and possible bias. If the target indicator is “was the candidate hired”, the model learns to imitate past selection rather than to predict future success. Even a retention indicator is ambiguous: departure may depend on the quality of management, pay, schedule and the external market, rather than on the employee’s “suitability”.
Therefore automation begins with verifying the target criterion. If the organisation cannot explain why the indicator is related to the requirements of the job and how alternative causes of the outcome are taken into account, the model merely gives an old decision new speed.
Proxy features and selection bias
A model can reconstruct sensitive information from indirect features: geography, educational institution, gaps in employment, linguistic constructions or the structure of a career path. Removing the “sex” or “age” field does not prove the absence of a discriminatory effect.
In addition, historical data contains only observed outcomes. The organisation knows the work results of hired candidates, but does not know how the rejected ones would have performed. This is selection bias: the sample of outcomes has already been formed by previous selection.
A systematic review of research on algorithmic decisions in recruitment and HR development shows that risk arises not only inside the algorithm. It can be created by task formulation, data, the interface, recruiter actions and the distribution of attention among candidates (systematic review). Therefore fairness cannot be checked with a single model metric.
Plausible generation instead of fact
A generative model can create a confidently written but unverified statement – NIST calls this risk confabulation (NIST Generative AI Profile). In HR this is especially dangerous when the model:
- attributes to a candidate experience they do not have;
- turns an ambiguous interview note into a categorical conclusion;
- misstates internal policy;
- adds reasons for a disciplinary decision that are absent from the source documents.
Therefore generative text must not become a fact of a personnel file on the basis of smooth wording alone. References to source records, verification of claims and a prohibition on adding unverified characteristics of a person are needed.
Drift and feedback loop
Drift – a change in data, conditions or the relationship between features and outcome after the model has been validated. It arises with a change in the labour market, role requirements, hiring channels, candidate composition, company policy or model version.
A feedback loop appears when the system’s decisions shape its future data. If a ranking more often shows recruiters candidates of a certain type, it is exactly those who more often pass the interview and enter the dataset of successful hires. The system then takes the pattern it created as confirmation of its own correctness.
A one-off acceptance before launch does not detect these effects. Monitoring, analysis of overrides (manual cancellation of the result) and appeals, version control and a pre-established ability to stop are needed.
The process is rebuilt first
Before choosing a model, the process owner must describe the chain from goal to consequences:
goal → input data → criterion → system output → human or system action → notification → review → outcome → monitoring
For each transition, four questions need an answer:
- What decision is actually being made?
- On what grounds is it permissible and related to the job?
- Who has the authority to change it?
- How can the person it affects report an error?
Then the baseline is recorded – the current state of the process without the new tool. This is not necessarily a single final figure. For document processing the baseline may include the share of errors, processing time and the number of returns. For screening, at a minimum, results by funnel stages, error types and differences between relevant groups are needed. Otherwise a reduction in time is easily taken for an increase in quality.
For exceptions a separate route is created. A non-standard career path, an incomplete document, a disability, the need for reasonable accommodation or a conflict of sources must not automatically turn into a low score. The system must be able to pass such a case to a competent employee without a negative conclusion about the person.
Data: minimisation is not blindness
Work with data must begin with purpose limitation: each field is collected for a defined purpose and is not reused merely because it is technically available. The data minimisation principle requires limiting personal data to the necessary volume; this reduces privacy risk but does not eliminate bias (ICO guidance on security and data minimisation).
For each source, the following should be recorded:
- origin and original purpose of collection;
- owner and access rules;
- completeness, currency and permissible values;
- transformations from the source record to the model feature;
- retention and deletion periods;
- versions of the dataset, rules and model used for a specific output.
The sequence of transformations is called data lineage. Without it, it is impossible to establish why the system produced a particular result and whether the error affected other decisions.
A real contradiction arises: a fairness audit may require protected attributes that cannot be used without justification in operational scoring. The solution is neither uncontrolled addition of these fields nor abandonment of the check. Audit data should be separated from the working loop, access and use restricted, a retention period set, and the legal basis confirmed separately.
If the processing is capable of creating a high risk for people’s rights, a DPIA – data protection impact assessment – applies, that is, a prior assessment of necessity, proportionality, risks and safeguards. A DPIA is not a one-off attachment to a procurement: it has to be revised when the purpose, data, vendor, model or consequences of processing change.
Control must cover the lifecycle
A reliable loop does not reduce to an accuracy test before launch. It includes several connected levels.
Before launch
The use case, prohibited ways of use, target group, model version, quality criterion and the cost of different errors must be recorded. For significant decisions, not only average indicators are checked, but also results for relevant subgroups:
- selection rate (the share of group members who passed a particular stage of the process);
- false positive and false negative rates;
- stability of thresholds and calibration;
- cases with missing or conflicting data;
- change in result under an irrelevant substitution of a sensitive attribute or a proxy related to it.
Selection rate – the share of group members who passed a particular stage of the process. For significant decisions a single average accuracy is not enough: selection rate, false positive and false negative rates, as well as other relevant metrics, should be analysed taking into account the type of error, the stage of the process, group size and applicable law.
The last check is often called counterfactual testing. It is useful for finding instability, but does not prove the fairness of the whole system: real groups may differ in data structure and in conditions that a simple substitution of a single field cannot reproduce.
There is no universal set of fairness metrics. The choice depends on the stage of the process, the type of decision, the position, the available data and applicable law. Small samples also do not allow confident conclusions to be drawn from apparently large percentage differences.
During operation
For each significant output, logs are needed that allow reconstruction of:
- the input data used and its version;
- the version of the model, prompt, rules and retrieval corpus;
- the result, the model’s confidence measure or the rule that triggered;
- the user’s action;
- the override and the documented reason;
- any subsequent complaint, correction or incident.
Logging should be designed together with privacy and retention, not added afterwards. An excessive log itself becomes a personnel store of elevated risk.
Monitoring must track data quality, changes in distributions, errors by group, the share of manual overrides, complaints and discrepancies between versions. The stopping threshold is set before the incident. Examples of conditions – loss of a mandatory data source, an unexplained change in selection rate, the inability to reconstruct the version of an output or the detection of systematically false statements about people.
The system must have a kill switch – an organisational and technical means of terminating the automated action – and a safe rollback to a validated process. Stopping is useless if HR cannot process cases manually afterwards.
The NIST AI RMF approach organises risk management around the functions Govern, Map, Measure and Manage; ISO/IEC 42001:2023 sets requirements for an AI management system and its continual improvement (NIST AI RMF Core, ISO/IEC 42001:2023). These frameworks help build governance, but certification of the organisation does not prove the correctness of a specific HR decision.
Human oversight must change the outcome, not decorate the interface
The reviewer must receive not only the model’s recommendation but also sufficient context for an independent decision. Otherwise automation bias arises – the tendency to trust an automatically proposed option, especially under time pressure.
Working human oversight requires:
- training on the purpose and limitations of the system;
- access to source data and the grounds for the output;
- time for a substantive review;
- authority to override the recommendation without sanctions;
- the same procedure for comparable cases;
- documenting the reason for agreement or override where justified by the risk;
- passing a complex case to a specialist;
- an appeal channel for the candidate or employee.
If one recruiter checks all the recommendations while another accepts the ranking without looking, the organisation is in fact applying different procedures to comparable candidates. This is a process defect, not an individual user error.
Responsibility cannot be transferred to the vendor
| Role | Primary responsibility |
|---|---|
| HR process owner | Purpose, criteria, permissible actions, exception route and the outcome of the process |
| Data owner | Origin, quality, access, retention periods and data lineage |
| Model owner or vendor relationship | Versions, limitations, documentation, testing, changes and vendor incidents |
| Reviewer or line manager | Substantive review of the specific case and documentation of the decision |
| Privacy, legal and DPO | Legal bases, DPIA, data subject rights and processing restrictions in the applicable jurisdiction |
| Security | Access, protection of integrations, secrets, logs and data transmission channels |
| Risk/compliance and employee representatives | Independent verification of controls and participation required by applicable rules or agreements |
| Executive accountable person | Acceptance of residual risk, resources for control and the decision to stop |
The vendor must provide sufficient information about versions, data, functions, limitations, validation and changes. If a closed SaaS system does not allow access to logs, version history, feature definitions and the results of relevant tests, the organisation cannot compensate for this gap with trust in the brand or a general presentation of accuracy.
Official UK guidance on responsible AI in recruitment recommends an assurance approach both to procurement and to the operation of systems at the sourcing, screening, interview and selection stages (Responsible AI in Recruitment). The practical meaning of this approach is to check not only the product demonstration, but also the organisation’s ability to obtain the evidence needed for its own control.
Example: document processing during onboarding
Consider an illustrative scenario. An organisation wants to reduce manual checking of the document package of a new employee.
It is safer to frame the task like this: the system extracts the document type, date and identifier, checks for mandatory fields and shows the operator the source fragment. It does not decide whether a person has the right to work, does not interpret medical information and does not create a reliability assessment.
The process can be built as follows:
- Completeness rules are defined outside the generative model and are versioned.
- Each extracted field is linked to a coordinate or fragment of the source document.
- Low confidence, conflicting documents and unknown formats are routed to the operator.
- Until the operator confirms, the details are not recorded as final in the HRIS (HR information system, i.e. a human resources information system).
- Errors are classified by document type and field, not merely averaged.
- A change of model or rules triggers re-validation on a fixed set of examples.
- Source documents and logs are stored for the established periods and are accessible only to the relevant roles.
Here AI speeds up a verifiable operation. The cost of error is limited by the fact that the system does not make a personnel decision and the operator sees the primary source.
If the same system is extended to output such as “the candidate looks suspicious” or to automatic termination of the process because of a mismatch, not only the function changes. A new criterion appears, a different risk, additional rights of the person and the need for a separate legal and technical assessment. This cannot be treated as a small update to the previous automation.
How to make a decision about deployment
The four categories below turn the class from the applicability map into an operational deployment decision, not a new classification.
Automate
Suitable for verifiable operations with limited consequences of error. An owner, access control, quality criterion, error log and manual handling route are needed.
Automate only with mandatory control
Suitable for draft generation, classification and recommendations that may affect a person. Verifiable sources, human oversight, logs, restrictions on further use and periodic monitoring are required.
Pilot with measurement
Suitable for ranking, matching and forecasts when the goal looks justified but local data has not yet confirmed quality and the absence of an unacceptable effect. A pilot must not covertly influence real decisions. Before it, the baseline, metrics, relevant groups, stop conditions and the criteria for moving to operation are defined.
Do not deploy
This covers prohibited practices, autonomous actions with unacceptable consequences, scenarios without a valid ground truth, systems without the necessary documentation and cases where the organisation cannot provide review, appeal or a stop.
The final decision should rest not on a list of model capabilities but on a package of evidence: the process map, purpose and criterion, data inventory, DPIA where it is needed, local validation results, roles and authority, monitoring rules, stop conditions and vendor commitments.
HR automation brings lasting benefit where the technology reduces mechanical work rather than dilutes responsibility. The more the output affects access to work, income, career or working conditions, the less basis there is for autonomy and the more requirements on the process, data, evidence and the person’s ability to obtain a review.