Skip to main content

Introducing the AI Inherent Risk Scale (AIIRS)

Mark A. Bassett, AI strategy, governance, and integrity in higher education, Associate Professor and Academic Lead for AI, Charles Sturt University, EDSAFE AI Catalyst Fellow.

With Kelly Webb-Davies and Ella Wicks.

Overview

 The AI Inherent Risk Scale (AIIRS) provides a structured approach for classifying tasks that use generative artificial intelligence (GenAI) into LOW, MEDIUM, or HIGH inherent-risk bands.

Classification is determined via three criteria—epistemic dependence, verifiability, and consequences of error—that define the nature and significance of a task’s reliance on GenAI. These criteria consider the extent to which GenAI is expected to supply information, the degree to which the output can be independently verified, and the seriousness of any potential errors.

AIIRS provides a consistent and defensible basis for assessing the inherent risk associated with GenAI-assisted tasks.

Purpose

The purpose of AIIRS is not to determine whether GenAI should be used, but to establish the level of inherent risk associated with a task that may require active management. AIIRS focuses on the inherent characteristics of a task, not on individual behaviour or user intent. Once a task’s inherent risk is understood, any additional safeguards, mitigations, or design choices may be applied where warranted, in line with any applicable governance arrangements.

AIIRS does not replace or override institutional policy, regulatory obligations, assessment design decisions, or the exercise of human judgement.

Scope

AIIRS is a classification instrument only, which indicates the level of risk that should be actively managed for a task that uses GenAI. It does not determine whether GenAI use is permitted, prohibited, ethical, compliant, or appropriate in any given context. AIIRS is designed for task-bounded human use of GenAI and does not cover autonomous or agentic AI systems, which introduce additional risks beyond the scope of this classification instrument.

Alignment

AIIRS does not replace or override institutional policy, regulatory obligations, assessment design decisions, or the exercise of human judgement. Classification outcomes must be interpreted and acted upon within existing governance, policy, and decision-making frameworks.

The Australian Higher Education Standards Framework (HESF) requires providers to identify risks to academic quality and integrity and to manage those risks through informed judgement and established governance processes. AIIRSsupports this requirement by providing a shared, task-focused method for classifying the inherent risk of GenAI use that can be applied by staff and students, while ensuring that decisions about safeguards, assessment design, and integrity responses remain within existing institutional governance, policy, and quality-assurance frameworks.

Classifications

Classification criteria

Epistemic dependence

Epistemic dependence captures whether a task requires the system’s representations of the world to be correct in order for the task outcome to be usable. Tasks with lower epistemic dependence rely only on user-provided material, without requiring the system’s representations of the world to be correct for the task outcome to be usable. Tasks with higher epistemic dependence require the system’s representations of the world to be correct for the task outcome to be usable.

Verifiability

Verifiability captures the basis on which the correctness of a GenAI system’s output can be verified for the task. Verifiability is assessed independently of consequences. A task may be high risk due to the requirement for expert verifiability, even where the immediate consequences of error are limited. Tasks with embedded verifiability enable quick, reliable verification by the user or the surrounding process, without requiring specific domain expertise. Tasks requiring expert verifiability depend on specialised expertise or external investigation that requires evaluative judgement.

Consequences of error

The consequences of error reflect the extent to which incorrect, misleading, or incomplete GenAI outputs affect decisions, records, or outcomes related to the task. Tasks with minimal consequences of error are those in which errors have minimal impact on understanding or outputs and do not affect decisions, records, or outcomes relating to people beyond the task. Tasks with significant consequences of error are those in which errors affect decisions about people, alter records relating to them, or compromise outputs that have consequences for individuals or groups beyond the task.

Classification model

AIIRS uses a max-dominant classification model that supports proportionate risk management by ensuring that any single high-risk feature of a task is not offset by lower-risk features elsewhere.

The AIIRS decision flowchart.

When a task is classified as HIGH risk

Tasks classified as HIGH must not proceed in their current form. One or more of the following interventions are required:

When a task is classified as MEDIUM risk

Tasks classified as MEDIUM require proportionate controls to manage identified risk. The following controls and conditions apply:

When a task is classified as LOW risk

Tasks classified as LOW require routine care appropriate to the task and context. The following routine practices apply:

Licensing

AIIRS is released under a Creative Commons Attribution–NonCommercial–ShareAlike 4.0 International (BY-NC-SA 4.0) license. Users may remix, transform, and build upon the work, provided they give appropriate attribution, do not use it for commercial purposes, and distribute any derivative works under the same CC BY-NC-SA 4.0 licence.

Download

Visit the official AIIRS website at http://aiirs.ai to download the slide deck.