When assistance becomes authority
Independent judgement and professional responsibility in AI assisted work
A person can agree with an AI recommendation, check every figure it contains and still fail to exercise independent judgement. The analysis may answer a narrowly defined question accurately while leaving the decision that actually matters unresolved. Human approval can then lend professional authority to assumptions no one has examined.
Imagine a partner considering whether to automate part of a professional service. An AI-assisted business case sets out the time saved and the projected reduction in delivery costs. She checks the calculations and accepts the recommendation. Yet the analysis gives little attention to where junior colleagues will acquire the expertise needed to review that work, or how the change affects the advice clients receive. The arithmetic withstands scrutiny while the investment case remains incomplete.
An organisation can become proficient at checking the internal consistency of an analysis whose assumptions it has never seriously considered. For firms whose clients buy judgement, that is a failure within the service itself. Professional responsibility includes examining how the problem was defined and whether the evidence can support the conclusion. Review that begins only when an answer is ready for approval may begin too late.
Philippe Aghion and Jean Tirole distinguish formal authority, the entitlement to decide, from real authority, effective control over a decision. Their theory explains how the distribution of information can separate the two. [1] Applied to AI-assisted work, this distinction directs attention to the process that determines what reaches the person signing off. Whoever shapes the account of the problem can influence the decision without holding the right to make it.
Delegation is necessary. Partners rely on associates, executives on analysts, and specialists on one another. The difficulty arises when the recipient cannot examine the basis of advice, or when the workflow accommodates only a narrow account of the problem. AI can contribute to that narrowing through the information it selects and the interpretation it supplies.
We should be precise about bias. AI outputs reflect the data and design choices behind a system, together with the instructions and information supplied in its use. Those choices can reproduce inequalities or make some considerations easier to recognise than others. NIST’s analysis encompasses institutional and societal conditions, statistical and computational processes, and human cognition. [2] A neutral tone provides no assurance that these influences have been examined. Human intuition offers no such assurance either. Every analysis necessarily selects; the selection must be justified against the purpose of the decision.
That scrutiny starts with deciding what to ask. In six months of fieldwork within a corporate data science team, Samir Passi and Solon Barocas examined how organisational aims become technical problems through choices about targets and proxies: what to predict, and what measurable quantity will represent what matters. Their work locates ethical scrutiny within that translation. [3] A system’s success against a chosen measure cannot settle whether the measure is appropriate.
My work with AMP Concept on driver, vehicle-cleaning and check-in operations makes this a practical question about what records allow management to conclude. Knowing how many hours an employee worked does not, by itself, establish how productively those hours were used. That assessment also requires reliable activity records and an understanding of the work available. A low output figure may warrant investigation without establishing individual underperformance.
The operating model therefore keeps recorded hours, vehicle movements and the context needed to interpret performance distinct. A company overview can help reconcile total hours while remaining insufficient for an individual productivity ranking. Missing activity durations remain incomplete records; they are not treated as zero. [4] The design makes some conclusions unavailable until the necessary evidence exists. This gives management a basis for investigating capacity without prematurely turning gaps in the record into findings about a person.
Now consider asking an AI system to explain supposed underperformance in records like these. The request would already concede a premise that needs examination. This is a hypothetical use of AI to interpret operational evidence; the AMP example illustrates the controls governing that evidence. The same discipline applies to professional advice. A documented fact, an inference from it and a recommendation each require justification. Review should make the reasoning between them visible, including the point at which the evidence runs out.
Expertise alone does not guarantee effective review. In a preregistered experiment involving 758 consultants, Fabrizio Dell’Acqua and colleagues found that GPT-4 improved performance on tasks within its capabilities. On a separate task outside those capabilities, consultants with AI access were about 19 percentage points less likely to reach the correct solution than those without it. [5] This concerns a particular model and experimental task. It shows why professional experience and success on neighbouring tasks cannot establish that a particular use of AI will improve decisions.
For consequential recommendations, I would make the test explicit: can the person approving the decision justify its objective, identify the evidence supporting it and explain why the conclusion follows? The record should show what they actually examined. Decision criteria set before reviewing the AI recommendation, documented source checks and the assessment of a credible alternative provide evidence of that work. Reviewers should identify any material uncertainty and what would lead them to reconsider. The depth of examination should reflect the stakes. Automatic disagreement would be as poor a substitute for reasoning as automatic agreement.
This requires access to underlying material, including relevant information absent from the summary. Asking the system for a more persuasive explanation does not establish independent support for its conclusion. A second AI tool may help locate weaknesses, but its contribution also needs examination against evidence. Independence concerns the basis on which a conclusion is justified.
Review arrangements need testing in practice. In a 2021 experiment involving 199 participants and simulated AI advice about food substitutions, Zana Buçinca and colleagues tested interventions including an initial decision before seeing the AI suggestion. Compared with simpler interfaces offering explanations, the interventions reduced reliance on incorrect advice when identifying the main carbohydrate source. Across all trials, overall task performance did not significantly exceed that of the simpler interfaces, and the interventions were rated more complex. [6] My inference is that requiring visible engagement may help, but organisations must establish whether it improves decisions in their own setting.
Management determines whether that engagement is feasible. An automation business case built on near-instant approval leaves little capacity to investigate an omission. Those who set workloads and handle exceptions share responsibility for the resulting decisions. Madeleine Clare Elish’s “moral crumple zone” describes how a person with limited control over an automated system can absorb responsibility for its failure. [7] Leaders should examine whether reviewers possess the information and authority their responsibility requires, including the ability to delay a decision when the evidence is inadequate.
The ability to exercise judgement also has to be developed. AI can contribute to that development. Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied a staggered rollout using data from 5,172 customer-support agents. They estimated an average 15 per cent increase in issues resolved per hour, with substantial benefits for less experienced workers. Shorter chats during system outages also provided evidence consistent with learning, although the authors caution that those estimates are noisy. [8] These findings concern one firm and occupation, so their implications for other kinds of work need examination.
For professional-services leaders, the question is whether their deployment develops the expertise it will later depend on. A better document produced with assistance does not establish that its author has become better at assessing a difficult problem. Junior colleagues need opportunities to develop and defend their reasoning, receive feedback and understand why a convincing answer fails. A firm relying on future expert reviewers has to provide a way for those experts to be formed.
The person affected by a decision also holds knowledge that review may need. An employee’s circumstances may be absent from the records; a client’s priorities may be poorly represented by the chosen objective. Respecting their agency requires a route through which relevant circumstances can be heard and change the assessment. Review remains incomplete when it overlooks the people whose situation the analysis purports to explain.
This connects with my doctoral work on social robots and the questions I am developing through The Uncanny Project about autonomy and human dignity. Care and childhood involve distinct vulnerabilities, but they sharpen a question relevant here: as assistance becomes more influential, can a person still introduce a concern the system has missed and have it change what happens? In professional work, that possibility should be designed into how decisions are made.
For Synesis, these requirements belong in the operating model. A consequential workflow needs an explicit decision purpose and a record of what evidence supports the proposed action. Relevant uncertainty must reach the person responsible for acting, who needs time and authority to seek further information. The practical test is then to examine decisions under realistic workloads: whether reviewers notice material omissions, question unsuitable measures and change course when stronger evidence warrants it. Cases in which they correctly accept the system’s advice belong in that evaluation too.
Approval and override rates have little meaning without this context. A high approval rate may reflect excellent assistance or shallow review. Frequent disagreement may reflect expertise, weak system performance or human error. The organisation needs to understand the reasoning and consequences behind those numbers before treating them as evidence that oversight works.
That standard also strengthens the commercial case. Time saved in preparing an answer should be assessed alongside the work of validation and correction, and the quality of the final decision. An apparent saving that depends on leaving material assumptions unexamined offers an incomplete account of value. A firm can justify wider adoption when it demonstrates better work with review proportionate to the consequences.
Independent judgement is part of what a professional firm is paid to provide. AI can extend the information available and help people reason more effectively; the firm remains responsible for what it concludes and recommends. Trust becomes justified when it can explain the basis of its advice, respond to contrary evidence and correct an assessment that no longer holds. A person’s agreement with an AI answer becomes meaningful when the organisation can show the judgement behind it.