A participant at a recent ASQA webinar asked a question that grows sharper as AI tools spread through the sector: if an RTO generates an assessment with AI after giving the model everything about the unit, then puts that assessment through validation, is the approach sound? The structure is permissible, because the Standards regulate the result rather than the method of creation. The real question is whether the tool that emerges is valid, consistent with the training product, and capable of producing authentic and sufficient evidence, and whether the humans reviewing it can reliably detect when it is not. This article sets out the legal tests at each stage, the specific ways AI-generated tools tend to fail them, and a pre-deployment checklist, with consequences for RTOs, assessment designers, validators and the students whose results depend on the tool being right.
A Well-Framed Question
The question is well-framed because it recognises that validation is not an alternative to good assessment design but a quality check on it, and it asks whether AI-generated design followed by human review through validation produces a compliant outcome. The answer is that the workflow is legally permissible in its structure, but its compliance depends entirely on the quality of the human review at two gates: the pre-use review before deployment, and the validation that follows. The question is not whether AI can generate an assessment tool, because it demonstrably can. The question is whether the tool it generates is valid, consistent with the training product, and capable of producing authentic and sufficient evidence of competency, and whether the reviewers can reliably detect when it is not. This article examines the tests that apply at each stage, identifies where AI-generated tools are most likely to fail them, and provides a checklist to reduce the risk of deploying a tool that validation later finds non-compliant.
1. The Workflow and the Legal Tests That Apply to It
The AI-generate-then-validate workflow has three stages. The first is generation: the RTO uses an AI tool to produce draft instruments after giving the model the unit's performance criteria, knowledge evidence, performance evidence and assessment conditions. The second is pre-use review: the RTO reviews the draft against Outcome Standard 1.3(b) before deploying it. The third is validation: the RTO subjects the deployed tool to the Outcome Standard 1.5 process, which confirms whether the tool is consistent with the training product and capable of producing valid, reliable judgements.
Each stage has a distinct test. Generation has no standalone Standards obligation: the instrument does not regulate how a tool is created, only whether the result satisfies the requirements. Pre-use review is governed by Outcome Standard 1.3(b), which requires every assessment tool to be reviewed prior to use to ensure assessment can be conducted in a way that is consistent with the principles of assessment and the rules of evidence set out under Outcome Standard 1.4. Validation is governed by Outcome Standard 1.5, which quality-assures the assessment system by validating assessment practices and judgements and confirming that the tools are consistent with the training product. Reviewing a tool before deployment and validating it afterwards are both required regardless of how the tool was created, so the structure the webinar participant described is permissible. Whether it reliably produces tools that pass both gates is the substantive question, and answering it requires understanding the specific ways AI-generated tools tend to fail, because a reviewer who does not know what to look for will not detect them.
|
The Three-Stage Workflow and Its Legal Tests |
|
Stage one, AI generation, carries no standalone Standards obligation: using AI to draft is equivalent to using any other drafting tool. Stage two, pre-use review, is governed by Outcome Standard 1.3(b): every tool must be reviewed before use to ensure assessment can be conducted consistently with the principles of assessment and rules of evidence. Stage three, validation, is governed by Outcome Standard 1.5: it confirms the tool is consistent with the training product and produces valid, reliable judgements. The obligations attach to the result, not the method. |
2. Outcome Standard 1.3(b): The Pre-Use Review and What It Demands of AI-Generated Tools
Outcome Standard 1.3(b) requires assessment tools to be reviewed prior to use to ensure assessment can be conducted consistently with the principles of assessment and rules of evidence under Outcome Standard 1.4, and 1.3(c) requires the outcomes of that review to inform any necessary changes. This is not a discretionary quality step. It is a mandatory precondition for deployment, and no tool, however it was created, may be used for summative assessment without satisfying it.
The principles of assessment under Outcome Standard 1.4 are fairness, flexibility, validity and reliability. The rules of evidence are validity, sufficiency, authenticity and currency. A pre-use review of an AI-generated tool must test whether the tool can produce evidence that satisfies all of these. The difficulty is that AI-generated tools tend to produce content that appears to satisfy the requirements on a superficial reading but fails on closer examination, and understanding why requires understanding how the content is produced. A large language model generates text by predicting sequences of words that are statistically consistent with patterns in its training data. Prompted with information about a unit and asked to generate assessment questions, it produces text that looks like assessment questions, uses the vocabulary of competency-based assessment, and references the performance criteria supplied. What it does not do is verify that the questions actually assess the specific competency described, that the evidence they would elicit would be sufficient for a competency judgement, or that the assessment conditions in the training product have been correctly interpreted.
This gap between appearance and function is the core risk. A reviewer who reads an AI-generated knowledge question and finds it clear, relevant-sounding and well-worded has not yet determined whether it measures the knowledge it claims to, whether the answer key is accurate, whether it reflects the actual knowledge evidence of the unit, or whether a student who answers it correctly has demonstrated what the training product requires. Those determinations require expertise the AI did not have when it generated the question.
|
The Tool Looks Right; That Is Not the Test |
|
An AI tool generates text that statistically resembles good assessment. It does not verify that the content actually measures the competency it names, that the answer key is correct, or that the assessment conditions have been honoured. The pre-use reviewer must supply the expertise the AI tool did not have, and must do it with the unit documentation open alongside the tool. A tool that reads well in isolation can still be wrong in every way that matters. |
3. The Specific Failure Modes of AI-Generated Assessment Tools
Knowing the specific failure modes gives reviewers and validators a focused inspection agenda. Five are observed most consistently, each mapped to the requirement it breaches.
The first is hallucinated requirements. AI models sometimes generate assessment conditions, knowledge evidence or performance criteria that do not exist in the actual unit, particularly for older or less commonly represented training products. The content sounds authoritative and uses correct formatting, but references things the unit does not specify, and a reviewer without the unit documentation open will not detect it. This breaches the requirement that assessment be consistent with the training product: a tool that tests requirements not in the product is not consistent with it, however well-structured it appears.
The second is incorrect answer keys. The model generates plausible answers to its own questions, but where an answer is technically complex or recently updated, the key may be wrong, outdated or misleadingly simplified, because the model does not have the current edition of the relevant standard, regulation or technical reference unless it was in its training data and has not since been superseded. The risk is highest in fast-changing areas such as aged care regulation, work health and safety, and environmental compliance. This breaches validity within the rules of evidence: a tool marked against an incorrect key does not produce evidence that the assessor can rely on to be assured of competency, and a student marked not yet competent because a correct answer diverged from a wrong key has been assessed unfairly, breaching the fairness principle.
The third is misalignment between the assessment method and the unit's assessment conditions. Many units specify conditions that constrain how assessment may be conducted: a real or simulated workplace, specified equipment, supervision, or a number of instances. AI models often default to written knowledge questions, which may not satisfy conditions requiring practical observation or supervised performance. A tool consisting entirely of written questions for a unit whose conditions specify workplace observation has not been built to the unit's requirements, which goes to validity: the method must be appropriate to the training product and its specified conditions.
The fourth is insufficient coverage of performance evidence. Performance evidence specifies what a student must demonstrate, including the number of instances, the range of contexts and the conditions of observation. AI models frequently address the performance criteria comprehensively while understating or omitting the performance evidence specifications, producing a tool that covers what the student must do but not how many times or across what range it must be observed. This breaches the sufficiency rule, which requires the quality, quantity and relevance of evidence to support an informed judgement of competency.
The fifth is context-free generic tasks. Where a unit is intended for a specific industry context, an AI prompted only with unit information may generate tasks that could apply to any sector, producing evidence of general knowledge rather than the industry-specific application the training product requires. Where the product is contextualised to an industry or work environment, the tools must reflect that, so a generic task that could appear in any qualification is again not consistent with the training product.
|
The Five Failure Modes at a Glance |
|
AI-generated tools most commonly fail in five ways: hallucinated requirements that do not exist in the training product; incorrect answer keys based on outdated or inaccurate information; misalignment between the assessment method and the unit's specified assessment conditions; insufficient coverage of the performance evidence requirements; and context-free generic tasks that ignore the product's industry context. Every pre-use review of an AI-generated tool must test specifically for all five, with the unit open alongside the tool. |
4. Outcome Standard 1.5: What Validation Must Confirm About AI-Generated Tools
Outcome Standard 1.5 quality-assures the assessment system, and when AI-generated tools are validated, the confirmation obligation applies in full. Validation does not grant a reduced standard of scrutiny because a tool was made with AI; it applies the same test it would apply to any other tool. The critical question is whether the validators have the expertise to detect the failure modes above. A validator checking formatting, coherence and face validity, but not closely cross-referencing the tool against the training product documentation, may miss hallucinated requirements, incorrect keys or misaligned conditions. These errors are not always visible to someone reading the tool in isolation; they are visible to someone reading it against the unit simultaneously.
This places a higher functional demand on validators than the validation of tools built by an experienced designer who worked directly from the training product. An experienced designer is unlikely to have invented requirements that do not exist; an AI prompted with product information may have done exactly that, and the validator must be equipped to tell the difference. The validation record for AI-generated tools should specifically address whether the tool was cross-referenced against the training product during validation, whether any discrepancies were found, whether the unit's assessment conditions are reflected in the tool's methods, whether the performance evidence requirements are fully represented, and whether answer key content was verified against current authoritative sources. A validation that confirms consistency with the training product without addressing these questions has not completed a validation appropriate to the tool's provenance.
The independence requirement adds a further constraint. Outcome Standard 1.5 requires that the outcome of a validation is not solely determined by a person who designed or delivered the training or assessment being validated. For an AI-generated tool, the person who crafted the prompts, supplied the training product information and selected which outputs to keep is the designer for this purpose. They cannot be the sole validator of the tool they prompted the AI to create. This matters because the common pattern is for one designer to generate a draft with AI, review and edit it personally, then submit it for validation; if that same person also drives the validation, the independence requirement is not met. At least one participant must be independent of the generation process and able to bring genuine critical scrutiny without investment in the design choices the AI output reflects. The point holds at scale too: an RTO that uses AI to generate tools across a whole qualification cluster and then validates them with the same team that generated them has not achieved meaningful independence, even if different members handled different units.
|
The Independence Requirement and AI Generation |
|
The person who crafted the prompts, supplied the product information and selected the AI outputs is the designer for the purposes of Outcome Standard 1.5. They cannot be the sole validator. At least one validator must be independent of the generation process and able to scrutinise the tool critically, with relevant industry expertise, without any investment in defending the design choices the AI produced. Self-validation of one's own prompts is not validation. |
5. Workflow Compliance Status: A Step-by-Step Assessment
The following table assesses each step in the workflow against the applicable obligation, distinguishing the compliant steps from the compliance risks that arise when a step is executed poorly.
|
Workflow step |
Status |
Standards analysis |
|
Designer prompts AI with training product information and generates draft tools |
Compliant |
No obligation governs the generation method. Using AI to draft is equivalent to any other drafting tool; the obligation attaches to the result |
|
Designer reviews the draft and selects content to retain, modify or discard |
Compliant |
This is the design function. Compliance depends on the quality of the review, not the use of AI. The designer must bring enough expertise to identify the five failure modes |
|
Designer submits the tool for pre-use review without cross-referencing it against the unit documentation |
Not compliant |
Outcome Standard 1.3(b) requires the review to confirm the tool enables assessment consistent with the principles and rules. A review that does not verify the tool against the unit cannot confirm this; hallucinated requirements and misaligned conditions go undetected |
|
Reviewer cross-references the tool systematically against the unit's performance criteria, knowledge evidence, performance evidence and assessment conditions, and documents discrepancies and their resolution |
Compliant |
This satisfies Outcome Standard 1.3(b). The review is specific, documented and connected to the training product, and the tool may be deployed once discrepancies are resolved |
|
RTO deploys the tool after pre-use review but never schedules it for validation within the five-year cycle |
Not compliant |
Outcome Standard 1.5 requires every training product to be validated at least every five years. Deployment is not a substitute for validation, and AI-generated tools must be in the schedule |
|
Validation is conducted solely by the person who wrote the prompts and designed the tool |
Not compliant |
Outcome Standard 1.5 requires that the outcome is not solely determined by the designer. The prompt author is the designer, so sole validation by that person fails the independence requirement |
|
Validation panel includes at least one person independent of the generation process, with relevant industry expertise, who cross-references the tool against the training product during the activity |
Compliant |
This satisfies both the collective-expertise and independence requirements of Outcome Standard 1.5. The record should document the cross-referencing and any discrepancies found |
|
Validation finds significant discrepancies, such as hallucinated requirements or incorrect keys, but the RTO keeps using the tool while the issues are addressed through an improvement action |
Not compliant |
Outcome Standard 1.5 requires validation outcomes to inform changes. A tool found inconsistent with the training product must be corrected before further summative use; continuing to use a tool with known validity failures breaches Outcome Standards 1.3 and 1.4 |
|
Validation finds the tool consistent with the product, a report documents the cross-referencing and panel composition, and the AI generation method is noted in the record |
Compliant |
A properly conducted, fully documented validation satisfies Outcome Standard 1.5. Noting the AI method creates transparency about provenance and supports future review cycles |
6. A Pre-Deployment Checklist for AI-Generated Assessment Tools
The pre-use review under Outcome Standard 1.3(b) is the critical quality gate before deployment. The following is the minimum inspection agenda for a compliant review of any AI-generated instrument, and every item must be verified with the unit documentation open alongside the tool.
|
# |
Checklist item |
Verification method and documentation |
|
1 |
Every performance criterion referenced in the tool exists in the current version of the unit |
Open the training product on training.gov.au and cross-check each reference. AI tools sometimes paraphrase, combine or invent performance criteria |
|
2 |
Every knowledge item tested exists in the knowledge evidence of the unit |
Match each knowledge question to a specific knowledge evidence item, and flag any that test knowledge not listed |
|
3 |
Performance evidence requirements are fully covered by the assessment methods |
Verify the tool's activities generate all required performance evidence, and count instances where the unit specifies a minimum number of observations |
|
4 |
Assessment methods are consistent with the unit's assessment conditions |
Read the assessment conditions: where the unit specifies workplace observation, the tool must include an observation component; written questions alone do not satisfy it |
|
5 |
All answer key content is verified against a current authoritative source |
For every question with a specific correct answer, check it against the current standard, legislation or technical reference, as AI may use outdated information |
|
6 |
Tasks are contextualised to the industry or workplace context of the qualification |
Check the packaging rules and context notes; scenarios and case studies should reference the specific industry rather than a generic setting |
|
7 |
The tool does not assess beyond the scope of the unit |
AI tools sometimes test competencies from adjacent units or a higher AQF level; verify all content is within the specific unit's scope |
|
8 |
The authenticity controls are appropriate for summative use |
Where the tool uses written responses, verify it includes mechanisms to support the authenticity rule; questions an AI can readily answer without student input need supplementary authenticity controls |
|
9 |
The tool has been reviewed by at least one person with current industry knowledge |
A technically accurate tool can still set tasks that do not reflect current practice; an industry-current reviewer should check the scenarios against real-world conditions |
|
10 |
Review findings are documented, discrepancies resolved, and the tool version and review date recorded before deployment |
Keep a review record with the reviewer's name and credentials, the date, the discrepancies found, the resolutions, and confirmation the tool is ready |
7. The Documentation Trail From Generation to Validated Deployment
The complete trail for a properly reviewed and validated AI-generated tool should contain five records in sequence, all available at audit. The generation record documents the AI tool used, the prompts provided, the date, and the version of the training product information supplied, establishing provenance so a future reviewer understands what the AI was working from. The pre-use review record documents the reviewer's name, credentials and role, the date, each checklist item with findings, any discrepancies and the changes made in response, and confirmation that the tool is ready, linked to the specific version that was deployed. The deployment and use record documents when the tool was first used for summative assessment, which cohorts were assessed, and any complaints, appeals or quality concerns arising, which informs the risk-based decision about when to validate. The validation record documents the product validated, the date, the panel composition with credentials and industry expertise, the tools reviewed, whether the tool was cross-referenced against the training product, the findings, and the resulting improvement actions, and it should note the AI generation method so future validators are alert to the specific risks. The improvement action record documents any deficiencies found, the changes made, their date, and any re-checking of student outcomes assessed using the pre-change version. Where validation finds an AI-generated tool contained hallucinated requirements or incorrect keys, the RTO must assess whether outcomes from the earlier version were affected and whether remediation is required.
Conclusion: Permissible in Structure, Earned in the Review
The workflow the webinar participant described is legitimate, but legitimacy is not automatic. AI can draft an assessment, and nothing in the Standards forbids it, but the instrument has never cared how a tool was made, only whether it works, and an AI tool produces something that looks like a working assessment without confirming that it is one. That places the entire compliance burden on the two human gates the workflow already contains: a pre-use review conducted with the unit open and the five failure modes in mind, and a validation that is genuinely independent and genuinely cross-referenced against the training product. Done that way, AI generation can save real time at the drafting stage while the human review supplies the assurance the model cannot. Done without that rigour, it produces tools that look compliant, deploy quietly, and surface their hallucinated requirements and wrong answer keys only when a validator, an auditor or an aggrieved student finally reads them against the unit. The tool the AI writes is a draft. The compliance is in the review.
|
Summary: Creating Assessments With AI and Validating Them |
|
1. Using AI to draft assessment tools is permissible: the Standards regulate the result, not the creation method, and the pre-use review and validation obligations apply in full. 2. AI-generated tools have five common failure modes: hallucinated requirements, incorrect answer keys, misaligned assessment conditions, insufficient performance evidence coverage, and context-free generic tasks. 3. All five require the reviewer to work with the unit documentation open alongside the tool, not to read the tool in isolation. 4. Outcome Standard 1.3(b) requires the pre-use review to confirm assessment can be conducted consistently with the principles of assessment and rules of evidence, through systematic cross-referencing against the training product. 5. Outcome Standard 1.5 applies the same validation test to AI-generated tools as to any other, but places a higher functional demand on validators to detect invented or inaccurate content. 6. The independence requirement means the person who wrote the prompts and selected the outputs is the designer and cannot be the sole validator. 7. A tool found in validation to be inconsistent with the training product must be corrected before further summative use; continuing to use it breaches Outcome Standards 1.3 and 1.4. 8. The pre-deployment checklist has ten items, each verified against the current unit on training.gov.au. 9. The documentation trail runs from a generation record through pre-use review, deployment, validation and improvement action, with the AI provenance noted in the validation record. 10. AI can write the draft, but the compliance is earned in the human review, not granted by the tool. |
References and Further Reading
Federal Register of Legislation (2025). National Vocational Education and Training Regulator (Outcome Standards for NVR Registered Training Organisations) Instrument 2025. https://www.legislation.gov.au
Australian Skills Quality Authority (2025). Practice Guide: Assessment. https://www.asqa.gov.au/how-we-regulate/revised-standards-rtos/practice-guides/practice-guide-assessment
Australian Skills Quality Authority (2025). Practice Guide: Validation. https://www.asqa.gov.au/how-we-regulate/revised-standards-rtos/practice-guides
Department of Employment and Workplace Relations (2025). Revised Standards for RTOs: Frequently Asked Questions. https://www.dewr.gov.au/standards-for-rtos/revised-standards-rtos-frequently-asked-questions
Australian Skills Quality Authority. National Register of VET (training.gov.au). https://training.gov.au



