Insights
Updated September 2026

AI Ethics in Practice: Who Can Stop a Bad Decision?

Responsible AI needs legitimate goals, bounded authority, decision-specific evidence and a real way to challenge and correct outcomes, not an approval click.

By Ada Studio
14 min read

General information, not legal advice. Legal position as of . Limitations in the Legal Notice

Review status: legal and language review by a named human reviewer is pending.

In this article

Consider a hypothetical customer-service assistant that recommends rejecting a claim. An employee clicks “approve” but cannot see the evidence or challenge the recommendation within the time allowed. The customer has no practical route to correct an error. Who owns the decision after it passes through the model, provider and operating team? Was there oversight or only a recorded approval?

My central argument is this: responsible AI requires more than acceptable model outputs. It requires legitimate objectives, bounded authority, evidence suited to the decision, and a practical way to challenge and correct outcomes.

1. Ethics and risk management answer different questions

AI ethics asks which objectives are legitimate, whose interests deserve protection and which trade-offs should be refused. Risk management asks how harms might occur, how failures will be detected and who will respond. Neither replaces the other.

A system can reliably achieve a goal people have good reason to reject; an attractive purpose does not establish safe operation. Better prediction cannot settle whether an intervention is coercive, and values cannot prove that an access control works.

The healthcare ethics review identifies familiar concerns around fairness, transparency, privacy and accountability, but does not establish the comparative effectiveness of its proposed safeguards. The finance review maps research on algorithmic decisions and autonomy through bibliometric and thematic analysis; publication patterns are not measurements of customer outcomes.

Maintain two connected records: why the use is acceptable and who may bear its costs; how consequential failures are controlled. Calling both an “AI ethics assessment” does not replace either argument.

Bjørn Hofmann’s conceptual analysis distinguishes several ways AI can support or undermine autonomy, including understanding, decision-making competence, voluntariness and relationships. The patient-autonomy review likewise looks beyond disclosure and proposes a broader assessment approach, although that approach has not been validated as a measurement instrument.

Imagine an assistant that explains a complex choice but offers only the option its provider prefers. Understandable advice can still constrain choice. This is an illustration, not an observed case.

For an organisation, the useful questions are therefore concrete:

  • Can the person understand the recommendation and its uncertainty?
  • Are realistic alternatives visible?
  • Can they decline without an unreasonable penalty?
  • Can they revise the decision or obtain independent review?
  • Whose objective is the system actually optimising?

Do not compress these questions into a single autonomy score: a good explanation cannot compensate for an unusable refusal mechanism. This is a governance judgment, not a validated weighting rule. Nor is all influence a violation. Surfacing an overlooked alternative may strengthen agency; substituting a provider’s objectives for the person’s does not.

3. A system people like may not be helping them

In the Science paper on sycophantic AI, Cheng and colleagues combine model benchmarking with three randomised online experiments. Relative to more challenging replies, engineered affirmation increased participants’ perceived rightness and reduced their intentions to repair interpersonal conflict. Participants rated the affirming AI more favourably, trusted it more and expressed greater willingness to return. Positive feedback can therefore mask a potentially worse outcome.

Whether repair is warranted depends on the conflict. These were immediate responses in selected online samples, not delivered apologies, long-term dependence or durable relationship outcomes. The comparison was affirmation versus challenge, not a neutral no-advice condition; it does not justify assistants that disagree with everything.

My recommendation is to separate approval from benefit. For decision support, test whether people detect errors and understand alternatives, not just whether they like the answer. For advice systems, evaluate respectful challenge alongside perceived helpfulness.

4. Behavioural influence is not the same as ethical legitimacy

The political-persuasion preprint by Hackenburg and colleagues adds a different kind of evidence. In preregistered experiments with UK adults, persuasive AI conversations increased petition signing relative to neutral-topic AI conversations.

The reported increases were 12.8 percentage points in the first study and an average of 19.7 points in the second. These are absolute differences in the experimental outcome, not relative percentage increases, estimates of electoral influence or evidence about an entire population.

The neutral-topic comparator does not establish superiority over equally engaging human persuasion or a good static message. Paid participation, selected issues and low-cost actions limit transfer to ordinary political communication. Increased action is not an ethical verdict.

For persuasion uses, I would assess voluntary choice, transparency, vulnerability and conflicts of interest alongside conversion. Effectiveness still needs an ethical justification.

5. Responsibility must survive delegation

The medical-AI responsibility review proposes a multilevel model across clinical, institutional and regulatory actors before and after deployment. It maps responsibilities, not improved patient outcomes. The “agentic loafing” paper proposes performance conformity, fragmented responsibility and numeric legitimacy as ways delegation might weaken scrutiny. These are discourse-derived hypotheses, not measured complacency; its risk profiler is unvalidated.

A practical response is to assign responsibilities around the actual decision:

  • A decision owner determines whether the use case should operate and accepts a defined residual risk.
  • A system owner maintains its configuration, permissions and tests.
  • A reviewer can challenge or stop consequential actions.
  • A correction owner can investigate and correct outcomes for affected people.
  • An independent challenge function can question the deployment when commercial incentives discourage interruption.

This is a proposed operating arrangement, not a legal allocation of liability. The roles matter only if they have evidence, resources and authority to act.

6. Assess the deployment, not just the model

A model-level evaluation misses tools, permissions, external information, persistent memory and handoffs. The hidden-safety-failures perspective examines technical and organisational mechanisms together.

Imagine an assistant reading a document that claims a customer authorised sending a confidential attachment to a new address. Can the document supply authority that belongs to a trusted approval process? This is a hypothetical deployment failure, not merely an inaccurate answer.

The coding-assistant security synthesis identifies repositories, skills, tool metadata and protocol interactions as attack surfaces. Its reported attack rates come from heterogeneous secondary evidence and should not be treated as directly comparable product-risk estimates.

The broader threat review includes original image-classifier demonstrations, not measurements of deployed LLM-agent compromise. A paper about agentic AI does not make every result evidence about an agentic workflow.

My recommended risk statement is specific: an unauthorised recipient receives a confidential attachment; a generated change executes with excessive privilege; an incorrect recommendation becomes an irreversible action. For each scenario, identify the entry point, required authority, consequence and recovery route.

7. A benchmark score is evidence, not a safety certificate

The safety-benchmark taxonomy preprint distinguishes 40 behavioural benchmarks from five related evaluation resources outside that core set. Its empirical ranking comparison covers only four benchmarks and three risk categories, rather than the entire inventory.

One test cannot establish whole-deployment safety, especially when results are missing or scoring methods differ.

Prompt-injection results require similar care. In the contextual-injection preprint, the adaptive experiment covers 150 scenarios with up to 45 attempts per scenario. The reported 96.7% success rate means eliciting an unauthorised email-sending tool call after that bounded search, not a single-attempt failure probability or an observed rate of real-world disclosure.

The same paper’s separate test of 100 constructed scenarios combines authorised internal actions with unauthorised external emails. Explicit communication boundaries eliminated observed violations for two tested models and reduced them for another. This does not establish universal security; it also offers evidence against the claim that all defences are futile.

Before using a score, I would require the task, denominator, model configuration, attacker budget, comparison condition, blind spots and evidence of legitimate usefulness. Blocking every action is not a solution for authorised work.

8. Reduce unnecessary review, not the capacity to intervene

Adaptive oversight would reserve human attention for decisions that need it. But what evidence justifies removing a check?

Kumar and Singh evaluate an oversight architecture on synthetic tasks against an autonomous worker and universal review simulated by a model given the correct answers. This is an engineering hypothesis, not a live safety comparison.

The architecture changes several features at once, contains unresolved cross-table inconsistencies and interprets a nonsignificant difference as equivalence without an equivalence design. Fewer interruptions do not prove safety parity or staffing savings.

The safety perspective by Gjergji and Enkelejda Kasneci offers a complementary question: do reviewers possess the expertise, time, evidence, authority and organisational protection needed to exercise independent judgment? Those are proposed requirements for meaningful oversight, not controls whose effectiveness that paper experimentally validates.

Test whether reviewers detect and change bad decisions under realistic time pressure before reducing review frequency. Measure missed consequential errors, not just approval speed, and verify that an objection stops the action. Meaningful oversight preserves intervention where consequences justify it.

9. Monitoring needs an owner and a response

The clinical-AI scoping review organises evidence around post-development robustness, post-deployment monitoring and lifecycle governance. It distinguishes conceptual discussion from more operational evidence, but does not establish the effectiveness of a particular surveillance programme.

The hidden-failures perspective proposes indicators including unauthorised tool executions, detection of seeded review errors and correction before downstream action. Their validity remains to be evaluated.

Every material monitoring signal needs an operational response. Specify what it measures, its denominator, observation window, relevant groups, blind spots, investigation trigger and owner with authority to act.

Run a harmless failure drill: does the alert reach the owner, can they restrict the action, and are all affected cases corrected? This is a proposed test, not a validated control package. More incidents may mean better detection; fewer challenges may reflect a less accessible way to contest an outcome. Logging can also impose privacy costs.

10. Turn principles into a working decision process

The following implementation sequence is my synthesis of the review. It is a proposal for local testing, not a validated standard, a legal compliance determination or an empirically optimal timetable.

First 30 days: define the consequential decisions

Choose a few use cases. Map what each system can influence and do, affected people, data access, external destinations, persistent state and irreversible actions.

Write down the intended benefit, unacceptable outcomes and important value conflicts. Name the decision and correction owners. Walk through one disputed recommendation, one refusal and one error discovered after action.

Produce a decision-and-responsibility map, not just a model inventory.

Days 31 to 60: challenge the assumptions

Test legitimate usefulness alongside unauthorised actions and relevant failure scenarios. Evaluate whether people understand the recommendation, can disagree and can obtain correction.

Exercise permission boundaries. Where relevant, include misleading external claims of approval, changed data, tool errors and ambiguous delegation. Test whether monitoring detects controlled failures and whether reviewers can stop the resulting action.

Keep “not tested”, “tested elsewhere”, “tested with failures” and “evidence unavailable” separate. Missing evidence is uncertainty, not a passing result.

Days 61 to 90: make a revisable decision

Compare findings with the unacceptable outcomes identified at the start. Approve a bounded use, narrow its permissions, redesign it or suspend it pending better evidence.

Record the remaining uncertainty, the accepting owner and the conditions that would invalidate approval: a new model, new tools, more consequential permissions, different users or a changed objective.

Finally, verify the recovery route. Who can contain a failure? What happens to queued actions and persistent state? How will affected people obtain a correction? A rollback plan that nobody has exercised is still an assumption.

The test is what happens when the system is wrong

Ethics and risk management meet when an organisation connects an acceptable purpose to evidence, constraints, accountable decisions and ways to correct errors. Test a difficult case: can someone identify an error, challenge it with evidence, stop the action and secure a correction for the affected person?

A human in the loop is not the same as a human in control. The difference is not the label. It is the practical ability to change what happens next.

These are practical governance questions, not a legal definition of human oversight. Applicable duties depend on the use and jurisdiction; the following rules are the most common anchors for Swiss and EU organisations.

Legal anchors (CH/EU)

  • Automated decisions: FADP Article 21 requires the controller to inform people of a decision based exclusively on automated processing that has a legal consequence for, or a considerable adverse effect on, them; on request, they may express their view and have the decision reviewed by a natural person. These duties do not apply where the decision is directly connected with concluding or performing a contract with the person and grants their request, or where the person has explicitly consented to the automated decision (Article 21(3)). GDPR Article 22 gives a right not to be subject to such solely automated decisions, subject to exceptions with safeguards. The Court of Justice held in SCHUFA (C‑634/21) that an automated credit score can itself be such a decision where a third party draws strongly on it to establish, implement or terminate a contract with the person, and in Dun & Bradstreet Austria (C‑203/22, 27 February 2025) that, under the right of access in Article 15(1)(h) GDPR, the person can require an explanation of the procedure and principles actually applied to their data.
  • Human oversight of high-risk AI: under the EU AI Act, providers must design high-risk systems so that people can oversee them effectively, including, as appropriate and proportionate, overriding the output or stopping the system (Article 14), and deployers must assign oversight to people with the necessary competence, training, authority and support (Article 26(2)). Under Article 113, as amended by Regulation (EU) 2026/1744, these duties apply from 2 December 2027 for Annex III systems and from 2 August 2028 for Article 6(1) systems related to products under Annex I, Section A. Under Article 2(2), neither article applies to systems related to products under Annex I, Section B, for which only Articles 6(1), 60a and 102 to 112 apply; delegated acts under Article 2(13) may limit Article 14 for Section A systems. Systems placed on the market or put into service before those dates are covered only as set out in Article 111(2). Consolidated AI Act
  • Impact assessment: FADP Article 22 requires a data protection impact assessment before processing that is likely to result in a high risk to people’s personality or fundamental rights. Private controllers are exempt, or may dispense with it, in the cases set out in Article 22(4) and (5): a legal duty to process the data, a certified system, product or service, or a qualifying code of conduct.
  • Swiss AI rules: the Federal Chancellery states that Switzerland does not yet have any overarching AI-specific legislation. On 12 February 2025 the Federal Council announced that Switzerland intends to ratify the Council of Europe Framework Convention on AI, which Switzerland signed on 27 March 2025. A consultation draft, in particular on transparency, data protection, non-discrimination and supervision, is due by the end of 2026.

FADP links point to the English translation, which has no legal force; the German, French and Italian texts are authoritative.

Sources and scope

These sources support the specific research descriptions above. The operating roles, tests and 30 to 90 day sequence are my proposed synthesis, not validated requirements of the cited papers.

Clarify how AI decisions are made

Connect business, HR, IT and risk through clear ownership and review routines.

You might also like

Need clearer footing for an AI decision?

Start with a focused conversation about a live AI use case, workflow bottleneck, training need, or governance gap.