← Learnbuilt

Authenticity and Currency Under Generative AI

A per-task audit of seven Australian VET assessment packages, tested against a compliance obligation that already exists

Author: Joshua Hubbard, Founder, Learnbuilt. TAE40116. Fifteen years across VET training, assessment and training management.
Published: July 2026
Contact: joshua@learnbuilt.com.au · learnbuilt.com.au

Abstract. Generative AI has not created a new compliance obligation for Australian registered training organisations (RTOs). It has exposed, at scale and at no cost, whether the assessment materials already in circulation meet an obligation that has been in the Standards the whole time. This paper reports a structured, element-by-element audit of seven complete VET assessment packages across three training package faculties, tested against the two Rules of Evidence that generative AI puts under the most pressure: authenticity and currency (Standards for RTOs 2025, s.1.4). The paper describes the two-stage, AI-assisted method used (Warrant and Ruling, disclosed in full in section 4.4: an AI language model drafts the calls, a qualified assessor reviews and owns them), reports the findings, states the limitations of the study plainly, and sets out seven options for RTOs, all of which sit inside a Standard the sector is already required to meet.


Executive summary

In March 2026 the Australian Skills Quality Authority (ASQA) placed artificial intelligence squarely inside the compliance conversation: its 3 March news item asked RTOs directly whether their use of AI is compliant with the 2025 Standards, and its 2026 Sector Workshops shared a set of draft AI Principles covering assessment quality and academic integrity (ASQA, 2026a). Sector commentary summarised the shift more bluntly, that AI in VET is now a standards issue and not merely an innovation one (CAQA Resources, 2026). Either way, the standard being pointed at already exists. Under Standard 1.4 of the Standards for RTOs 2025, authenticity (the evidence is the student’s own work) and currency (the evidence demonstrates current skills and knowledge) are two of the four Rules of Evidence every assessment judgement must satisfy. Generative AI does not change that requirement. It changes how easily an assessment can appear to meet it while failing to.

To see what a systematic test of that obligation reveals, this paper audits seven complete assessment packages, unit by unit and task by task, across business, community services and information technology. The corpus is a single set of productised builds constructed to a common assessment template (see Limitations, section 7); the value here is in the method and in what it surfaces, not in claiming sector-wide prevalence. Each package was first researched at the element level to establish where the underlying occupational task now legitimately uses AI and where the skill must still be shown unaided. Each assessment task was then ruled against two axes: is the evidence the student’s own, and does the task assess the work the way it is now done.

The findings are consistent.

None of these fixes requires new regulation. They are what Standard 1.4 already asks for, applied task by task instead of unit by unit. The paper closes with seven recommendations and, in the appendix, the full worked audit for all seven units with sources.


1. Introduction: the problem, framed where it belongs

An assessment task can map cleanly to every performance criterion, read as fully compliant, and prove almost nothing. That was true before generative AI. What AI changed is the price of producing that gap. A tool that writes a plausible workplace report, a plausible reflection, or a plausible ethical-dilemma response in seconds has made the take-home written artefact, long the default unit of VET assessment evidence, far weaker as proof that a student can do the work.

The two common responses are both wrong. The first is to ban AI and rely on detection. Text detectors are unreliable and biased against writers for whom English is a second language (Liang et al., 2023), and OpenAI withdrew its own AI-text classifier for low accuracy; a policy whose only enforcement is the hope of noticing is not a control, and the case that assessment design, not detection, is the defensible response is by now well made in the literature (Dawson et al., 2024). The second is to reach for the newest visible feature, an AI chatbot conducting a roleplay assessment, for example, because it looks like a serious answer to a serious problem. Some of these tools may have a legitimate place. But adopting one because it is impressive, rather than because an analysis of the specific competency showed it was the right evidence-gathering method, is the same mistake in the opposite direction: motion without diagnosis.

Between the ban and the shiny object sits the actual work: for each part of each unit, decide whether the skill must still be shown without help or whether using AI well is now part of the job, then build an assessment method that produces defensible evidence either way. That decision cannot be made at the level of “our RTO’s AI policy.” It has to be made element by element, because a single unit routinely contains both kinds of skill at once.

This paper reports what that element-by-element analysis found when applied to seven assessment packages. It is deliberately empirical. Units are named by their national codes, which are public. No RTO is named, and the corpus is described honestly in section 7: these are not a random sample of sector materials but a single provider’s productised builds, used to demonstrate the method and to surface the fault patterns it detects.

2. Who is actually responsible

Before the findings, a fair question: if these faults are this consistent, whose job is it to prevent them? Accountability splits three ways, and none of the three has full ownership. Naming that split is part of the argument.

The Jobs and Skills Councils own the content of the training packages. Future Skills Organisation, for example, maintains the BSB, FNS and ICT packages, and is the body with formal authority to write AI-augmented practice into a unit’s Elements, Performance Criteria and Knowledge Evidence. But units of competency are deliberately written at a durable, outcome level (“apply critical thinking to work practices”; “write complex documents”) precisely so they do not need rewriting every time a tool changes. That convention is good design in general. It is also the convention now leaving a genuine gap, because generative AI is a larger shift than the convention was built to absorb, and training package review runs on a multi-year cadence. To be fair, the fix is already moving at this level: Future Skills Organisation has live projects adding applied AI content into the ICT training package and expanding BSB to treat generalist AI as a core digital skill, with outcomes expected through 2026 (Future Skills Organisation, 2025). The gap is real, but it is not being ignored.

The regulator, ASQA, governs whether an RTO’s assessment system meets the Rules of Evidence, not how it does so. That flexibility is a feature. It lets one unit be assessed legitimately by a metropolitan RTO, a regional RTO and an online-only RTO without a single prescribed method forced on all three. It also means ASQA prescribing a specific AI-handling method would sit slightly outside its normal role, even after it placed AI inside the 2025-Standards compliance conversation in March 2026 (ASQA, 2026a). ASQA can say the obligation applies; it is not built to hand every provider the redesign.

Which leaves the individual RTO and its assessment designers. They are closest to the work, and least resourced to solve a fast-moving, cross-industry problem unit by unit with no shared reference point. That is the shape of the gap. It is not that anyone is asleep; it is that the system’s own sensible design pushes this particular decision down to the level with the least capacity to make it well at scale.

One data point sharpens the picture. The most visible market response so far is a privately owned accredited qualification, 11287NAT Diploma of Artificial Intelligence, licensed to RTOs to add to scope (training.gov.au). A nationally endorsed AI credential also exists inside the ICT training package with no licence fee (ICTSS00120 Artificial Intelligence Skill Set), though the two aim at different learners: the skill set at technical developers, the diploma at generalist business users. They do not cover the same need, which is itself the point. A fast-moving demand for AI capability is being met partly by a technical endorsed product, partly by a generalist private one, and unevenly overall. Internationally the OECD has published principles for using AI in VET (OECD, 2026), but those address building curriculum, not assessing students, so they do not close this gap either. Section 5.7 returns to how even the AI-specific qualification assesses its own subject.

3. The regulatory hook: already binding, and the regulator has said so

The weak framing for a paper like this is “the regulator has not caught up.” That framing is unnecessary and, as of 2026, untrue. In 2026 ASQA announced a Sector Workshop series, delivered over March and April, asking RTOs directly whether their use of AI is compliant with the 2025 Standards, and sharing a set of draft AI Principles built around assessment quality, academic integrity and the integrity risks AI introduces (ASQA, 2026a). The regulator has, in other words, already placed AI inside the standards conversation; sector commentators have put it more bluntly still, that AI in VET has moved from an innovation issue to a standards issue (CAQA Resources, 2026). Practitioner bodies are running professional development on exactly this question (VELG Training, 2025). What has been missing, to the author’s knowledge, is not attention but a worked, per-task method applied to real materials. (A separate document, ASQA’s AI Transparency Statement, governs ASQA’s own internal use of AI as a regulator, and should not be confused with guidance to providers (ASQA, 2026b).)

What ASQA has not yet published is the finished standard against which “compliant” will be judged; the draft Principles were presented at workshops, not released in full. But the underlying obligation does not depend on that document, because it already exists and pre-dates the March announcement. Under Standard 1.4 of the Standards for RTOs 2025, assessors must make judgements justified against four Rules of Evidence:

“(i) validity … (ii) sufficiency … (iii) authenticity, the assessor is assured that a VET student’s assessment evidence is the original and genuine work of that VET student; and (iv) currency, the assessment evidence presented to the assessor documents and demonstrates the VET student’s current skills and knowledge.” (Standards for RTOs 2025, s.1.4)

And Standard 1.3 requires that “assessment tools are reviewed prior to use to ensure assessment can be conducted in a way that is consistent with the principles of assessment and rules of evidence set out under Standard 1.4” (Standards for RTOs 2025, s.1.3).

Read together, the position is clear. Authenticity is precisely the property generative AI strains: is this the student’s own work, or the tool’s? Currency is the mirror property: does the task still assess the job as it is actually done, or a pre-AI version the workplace has left behind?

Two clarifications keep this honest, because “currency” carries more than one reading. In the Rules of Evidence it is most often read narrowly, as the recency of the student’s own evidence: how long ago the competency was demonstrated. This paper uses it in the wider sense the instrument’s own words allow, that the evidence must demonstrate the student’s current skills and knowledge, on the argument that if the workplace skill itself has changed to be AI-augmented, then evidence of the superseded, pre-AI version is no longer evidence of a current skill. A regulator-minded reader may not accept that extension, so nothing in this paper rests on it alone. Every finding routed to “currency” is independently a validity problem, the evidence no longer assures the skills as they are performed to the standard required in the workplace, and a Standard 1.3 problem, since a tool that assesses an obsolete workflow is not one that reflects contemporary industry practice on the pre-use review that Standard requires. A reader who rejects the extended reading of currency arrives at the same fix by the validity and 1.3 routes. The axis is a convenient label for where the pressure lands; the obligation behind it is over-determined, not fragile.

Every finding in this paper is therefore a failure against requirements that are already binding, made visible by AI rather than created by it. The method treats authenticity and currency as two axes to test, task by task, exactly as Standard 1.4 already frames them, and exactly what ASQA’s March 2026 position gestures toward without yet spelling out. It invents no new criterion, and gives an RTO a way to act now, ahead of the final guidance.

4. Method: Warrant and Ruling

The audit runs as two named stages, kept deliberately separate because they answer different questions and draw on different authority. The first, Warrant, is a research process: it establishes what the assessment should be testing given how the work has changed. The second, Ruling, is an audit process: it takes a real, as-built assessment and rules on whether it holds. Warrant investigates the unit; Ruling rules on the task. The full flow is shown in Figure 1.

Figure 1. The two-stage method. The unit of competency is fixed. Warrant researches how AI is reshaping the occupational task now and outputs a Resist / Integrate / Split call for each element; Ruling takes each real assessment task, scores it on two axes (authenticity, currency), and returns one of five decisions. What must be demonstrated never changes; only how the evidence is gathered.

4.1 Warrant: establishing the Resist/Integrate split

Warrant researches the underlying occupational task against how AI is actually reshaping that work in Australia now, weighed against a ranked source hierarchy: the regulators and the unit record first (training.gov.au, ASQA, TEQSA), then national research (NCVER, peer-reviewed assessment scholarship), then the industry skills bodies, down through sector peak bodies and named practitioners, with vendor and AI-detector marketing lowest and never counted as evidence. The research is deliberately anchored to Australian sources and to the unit’s actual occupational task, not the broad field; an overseas source is used only as a flagged exception where it is a globally authoritative primary study with no Australian equivalent. The output is an element-by-element call (Figure 2). Each part of the unit is marked:

Warrant never touches what must be demonstrated, which is fixed by the training package. It rules only on how the evidence should be gathered. Its research is documented and sourced per unit; those sources are listed in Appendix A.

This element-level call is close in spirit to the AI Assessment Scale (Perkins et al., 2023), the framework behind much Australian tertiary redesign, and differs from it in three ways that matter for VET. It is made per element of a unit, not per whole task; it is anchored to the VET Rules of Evidence and a unit’s fixed Assessment Conditions rather than to general higher-education pedagogy; and it is driven by research into the specific occupational task rather than chosen by the educator from a scale. The scale describes how much AI a task is permitted; Warrant researches how much the job has already changed.

Figure 2. The Resist/Integrate spectrum. Each element of a unit is placed from Resist (show it unaided) through Split to Integrate (AI-augmented performance is the skill), based on the research, not on a single unit-wide policy setting.

4.2 Ruling: auditing the real task on two axes

Ruling takes the as-built assessment at face value and scores each task on two axes (Figure 3):

Each task then resolves to exactly one decision, with no hedging:

Decision When it applies
Compliant Passes both axes; a verification mechanism exists
Revise A specific, named fix preserves coverage and salvages the method
Rebuild The method is the wrong evidence type for what the unit requires; a patch cannot fix it
Flag A decision only a human can make (a whole-of-package redesign; a coverage gap needing unit interpretation)
Escalate A missing input could change the call; the decision is still made under a stated assumption
Figure 3. The two-axis routing. Authenticity (does the evidence hold as the student’s own?) runs left to right; currency (does the task assess the work as it is now done?) runs bottom to top. A task that holds both is Compliant; authentic but obsolete routes to Revise (modernise); a task where own work is not assured routes to Revise or Rebuild. Package-level checks then apply: at least one verified task must exist, and a systemic pattern is named once, not patched task by task.

Two disciplines matter as much as the decisions. The method must be willing to leave a sound task alone: forcing AI into a unit where the work is genuinely AI-thin manufactures relevance and weakens validity. And it must distinguish a systemic fault (one design decision repeated across a whole package) from a cluster of unrelated faults that happen to co-occur, because the two call for very different fixes.

At summary level, the rules a second assessor would apply to reproduce a ruling are these. Run the outsourcing test on every task. A task a generative model could complete unsupervised, where the skill must be shown unaided, fails authenticity and is at least a Revise. A written task standing in for a required performance is a Rebuild, not a Revise, because the evidence type is wrong. Currency is scored only against the Resist/Integrate split, never re-derived, and an Integrate element assessed through its pre-AI workflow is a Revise (modernise). At package level, at least one verified task must exist or the package is Flagged, and where most tasks fail the same way the pattern is named once rather than patched task by task. The calibrated thresholds and the full decision tree behind these rules are the method’s detailed layer and are not reproduced here, but these summary rules are enough to re-run the two axes on the same units from their public requirements.

4.3 The corpus

The seven units were selected for faculty spread: four from business (BSBCRT411, BSBPEF201, BSBWHS411, BSBWRT411), two from community services (CHCDIV001, CHCLEG001), and one from information technology (ICTICT451). BSB and CHC are among the highest-enrolment training package areas nationally (NCVER, 2025), so the weighting reflects the sector. In the interest of disclosure, every package audited is one of the author’s own commercial builds, constructed to a single common assessment template; the faults reported below are faults in the author’s own prior work. That origin is a material limitation, addressed in section 7. The full worked audit for each unit, with sources, is in Appendix A. This study addresses the assessment of enrolled learners; recognition of prior learning and credit transfer, where the authenticity of a candidate’s evidence portfolio is arguably even more AI-exposed, are out of scope here.

4.4 The tools are AI-assisted, and the human is accountable

Warrant and Ruling are not metaphors for a manual process; they are AI tools. Each is a structured Claude project: a folder of instructions and reference files that an AI language model reads and then operates under. In Warrant, the AI performs the live research into how AI is reshaping the occupational task and drafts the Resist/Integrate call for each element. In Ruling, the AI applies the documented decision criteria to each assessment task and drafts the ruling. The human practitioner, a TAE-qualified assessor, configures the tools, directs each run, reviews every Resist/Integrate call and every ruling, and is accountable for all of them. The AI drafts; the human decides and owns the decision.

This is deliberate, and it is the same pattern the paper recommends for assessment itself (section 5.5): the AI does the research and the first pass, and the assessable, accountable act is the human judgement over that output. The audit is, in that sense, a worked example of its own recommendation rather than an exception to it. Two consequences follow and are stated plainly. First, where section 7 describes this as single-assessor expert judgement, that judgement is AI-assisted: a second person re-running the tools would still be reviewing AI-drafted calls, so an inter-rater check tests the human review layer, not a fully manual process. Second, the data-handling caution in section 6.7 applies to this audit too, because the assessment packages were fed into an AI system. Here the exposure is low, since the packages are the author’s own commercial products built around a fictional organisation and contain no real student or personal data, but the principle is the one the paper asks RTOs to observe.

5. Findings

The shape of the results across the seven packages is summarised below; the full per-task audits are in Appendix A. Because a package-level Flag or a systemic Revise deliberately collapses many failing tasks into one finding, the table reports the headline outcome per unit rather than a task count that would misrepresent that collapse.

Unit (faculty) Tasks audited Headline outcome Live checkpoint as built
BSBWHS411 (Business) 6 + policy Mixed: one Compliant, the rest Revise (both axes fire) Yes (observed consultation)
CHCDIV001 (Community services) 3 + policy Mostly Compliant; one Rebuild (wrong evidence type) No (AI-thin unit)
BSBCRT411 (Business) 19 + policy Package Flag (17 of 19 fail); one Compliant checkpoint and its compensated draft carved out Yes (recorded presentation)
BSBPEF201 (Business) 9 + policy Knowledge questions Flagged; project Compliant via a live role-play; two Revise (resources record and policy) Yes (assessor-graded role-play)
BSBWRT411 (Business) 6 clusters + policy Package Flag (every cluster fails; no verification anywhere) No
CHCLEG001 (Community services) 17 questions + project Systemic Revise; Rebuild on the three situation-response steps No
ICTICT451 (IT) 15 items Near-universal Revise (cluster-specific, not systemic); Flag for zero verification No

Escalate (defined in section 4.2) is not a standalone outcome here; it functions as a qualifier on another decision when a needed input is missing, and it applied once, to BSBCRT411, where the delivery mode was unstated and the call was made under the harder online-and-unmonitored assumption. No unit resolved to a standalone Escalate.

5.1 Every package carried the same AI-use policy fault, and the policy is not the point

Every package contained an acceptable-use statement of the same shape: AI may be used “to support your learning (drafting, summarising, checking spelling and grammar),” but AI-generated text may not be submitted as the student’s own work, with breaches routed to the RTO’s academic integrity policy. In this corpus the wording was identical because the packages share one template, which is exactly why this finding needs framing carefully rather than as a headline count.

In the author’s experience across RTOs, most already have a statement like this, in assessment cover sheets or the student handbook; having one has become common practice. That is not the fault. The fault is that a statement of this kind is a declaration, not a control. It tells the student what not to do and names a consequence if caught, but it changes nothing about whether a given task can actually be outsourced, and detection is not a mechanism (Dawson et al., 2024). An acceptable-use statement sitting over a set of take-home written tasks with no verification is doing no work at all. The fix is to make the policy operational: a two-tier approach that requires AI, with the prompt logged, on the specific tasks where it belongs, and prohibits it, with a disclosure requirement, everywhere else, so that the policy and the task design finally agree with each other. In three of the seven packages a blanket ban would directly contradict the modernised tasks the audit recommends, so the policy rewrite is not optional once any other fix lands.

5.2 Most packages had no verification anywhere, and that was a design decision, not an oversight

Four of the seven packages (CHCDIV001, BSBWRT411, CHCLEG001 and ICTICT451) had no live, oral or observed checkpoint anywhere in the assessment as built, despite covering competencies where a supervisor, client or colleague would ordinarily be present. In the starkest of them the package was built entirely take-home and entirely written, on the reasonable-sounding assumption that this is cheaper to deliver at scale, with no allowance for what that costs once a text generator exists. (Two packages failed uniformly enough to be resolved as a single finding rather than task-by-task; that grouping, discussed in section 5.6, is not the same as this one, because one of those two, BSBCRT411, does carry a live checkpoint.)

The packages that held up were the ones with a genuine live component already built in: an observed WHS consultation, an assessor-graded wellbeing conversation, a recorded stakeholder presentation. In each case the single live checkpoint did more than pass on its own; it lifted the written tasks feeding into it. To be precise about what this buys, an oral defence establishes that the student understands and can work with the submitted artefact, not, on its own, that they authored every line of it: a well-prepared student could defend an AI-drafted plan. What it does is convert a bare artefact into triangulated evidence, the product plus a live demonstration of command over it, which is usually what the competency actually requires. So one well-placed verification point can move several otherwise outsourceable tasks around it from “unverifiable” to “anchored.” This is the most actionable, lowest-cost, highest-value fix the audit surfaced, and it requires no AI policy at all. It is an assessment-design fix that would have mattered before generative AI existed; AI simply raised the cost of not having it.

5.3 The two faults behind most individual problems

Two faults accounted for the majority of task-level problems. The first is the scenario handed to the student. A named worker, a supplied case brief, a provided source pack: the facts the task claims to test are already in the prompt, so the task can be completed without those facts passing through the student’s own judgement. This appeared in every package in some form.

The second is a written response substituting for a demonstrated skill. Where a unit’s Assessment Conditions call for interaction, observation or “modelling of industry operating conditions,” a written “how I would respond” template is not a weaker version of the right evidence; it is the wrong evidence type. This is a structural fault, not a wording fault, and it recurred in the interpersonal, judgement-heavy community services units, where responding to a real person across difference, or navigating an ethical dilemma in the moment, cannot be evidenced by a description of what one would hypothetically do. Both such cases required a rebuild of the method, not a revision.

5.4 Not every unit needed changing, and the method has to be able to say so

One unit, CHCDIV001 (Work with diverse people), came through the audit almost entirely clean. Its core competency, reflecting on one’s own cultural bias and responding respectfully to real people across difference, is not meaningfully something AI can perform on a student’s behalf. Once a single evidence-type fix was applied, its tasks were sound as built. A method that can only ever recommend more AI-proofing is as unreliable as one that never catches anything. A second unit (BSBPEF201) showed the same restraint across three of its four elements. Across the seven units, the majority of assessed elements were left unmodernised, because the underlying work either does not meaningfully involve AI or must be shown unaided to protect authenticity. (This is a rough count offered as a description of this corpus, and the units are split at differing levels of granularity, so treat it as a direction of travel, not a measured rate.) This was not a “put AI in everything” exercise; where the audit did recommend integration, it was because Warrant found the real occupational task had already moved.

5.5 Where a unit genuinely is AI-augmented, the fix is judgement over output

Where a task assessed work that is now routinely AI-assisted in the real job (drafting a policy recommendation, summarising legislation, producing a first-draft report, analysing incident data), the fix took the same shape every time: require the AI-generated draft, capture the prompt used, and assess the student’s ability to find and correct at least two problems in it, then justify the corrections in their own terms. The competency shifts from producing the document to judging it. The student’s judgement over the tool’s output, not the output itself, becomes the evidence (consistent with the direction of TEQSA, 2025, in higher education).

The pattern generalised across unrelated subjects: legislation lookup in a community services unit, policy drafting in an IT unit, WHS communications in a business unit and commercial writing in another all converged on the identical fix. That convergence is a sign the fix is structural rather than subject-specific, which is what makes it teachable and repeatable.

One caution the audit built in: attaching this judgement task to an unsupervised take-home does not, on its own, close the authenticity gap, because a student could have the AI produce the “prompt,” the “problems found” and the “corrections” too. Modernising a task for currency can reopen the authenticity hole it was meant to help with. So wherever the fix was applied, authenticity was re-tested afterward and, where needed, a minimum verification anchor attached, a short oral walk-through of the corrections being the cheapest. Currency and authenticity are two axes, and a fix on one is not allowed to quietly break the other.

5.6 Some packages are broken as a system, not as a list

Two packages were resolved to a single package-level finding rather than task-by-task patches. BSBWRT411 failed on every one of its six task clusters, with no verification anywhere in the package. BSBCRT411 failed on seventeen of its nineteen tasks for the same reason, with one live checkpoint (a recorded presentation) and its compensated draft carved out as the two exceptions; the seventeen were named once as a systemic finding rather than patched individually. In both cases the correct call was to recommend a whole-of-instrument redesign around a single, centrally designed verification point, rather than hand back six or seventeen separate fixes that still do not add up to a defensible package. The size of the problem is not always the number of broken tasks; sometimes it is one design decision, repeated.

The IT unit showed the opposite discipline, and this is why the distinction matters. Nearly every one of its tasks needed a fix, but the fixes were genuinely different task to task: a judgement-over-output rewrite here, a verification anchor there, a split fix on a task that was half current and half obsolete. Collapsing those into one “systemic” finding would have hidden more than it revealed. Knowing the difference between one repeated fault and many different faults that co-occur is itself a finding. A sector response that cannot tell them apart will either over-flag units that are basically fine or under-fix units that need a genuine rebuild.

5.7 The qualification that teaches AI shows no public evidence of assessing for it

A closing observation, offered carefully. The most visible market response to AI in VET is 11287NAT Diploma of Artificial Intelligence, a privately owned accredited course delivered under licence by several RTOs. Its published assessment description is the conventional model, “answering knowledge questions and completing workplace-based practical assessment tasks,” with videoed roleplays where required (training.gov.au), and nothing visible in that public description addresses AI-specific integrity, on a qualification whose subject is AI itself. This was not among the seven packages audited, and whether the underlying instruments do more than the public description suggests cannot be determined from the register. The point is narrower than a verdict on the product: even here, in the corner of the sector most explicitly focused on AI, the public evidence that the assessment-design question has been asked is absent.

6. A menu of assessment methods, and cheaper alterations to existing ones

The fixes in this paper are not exotic, and the method does not invent new assessment types out of nothing. It draws on a documented menu of assessment-design patterns and fits them to the competency, rather than reaching for one favourite instrument. This section sets that menu out in general terms, independent of any single unit, as options an RTO can hold in mind. None of these are proprietary; they are the working vocabulary of good assessment design, made explicit.

The gate in front of every one of them is the outsourcing test: for any task, could a student run the whole thing through a generative model, unsupervised, and get the answer? If yes, and the skill must be shown unaided, the task is not defensible on its own and needs either controlled conditions or a verification layer. Most artefact-only take-home tasks fail this test, which is why so much of the audit turns on it.

6.1 Methods that protect authenticity (keep AI out where it must stay out)

For elements marked Resist, the goal is evidence AI cannot produce on the student’s behalf.

6.2 Methods that put AI in the open (assess judgement over the tool)

For elements marked Integrate, the assessable act is the student’s judgement over an AI output, not the unaided production of an artefact.

6.3 Verification patterns (establish authenticity without gold-plating)

Verification is the integrity anchor, applied minimum-sufficient, not everywhere.

One caution on the cheapest online-native option. An asynchronous recorded response is not a free pass: it can be coached in real time by a second screen or a voice assistant, so it partially fails the same outsourcing test that governs everything else. It is the right default for cost and fairness, but it is a floor, not a guarantee. Where the competency turns on unaided, in-the-moment judgement, the async tier has to escalate to a live, synchronous checkpoint with an unseen element, and the “does this fix reopen a hole?” discipline from section 5.5 applies to verification design just as it applies to task design.

6.4 Cheaper alterations to existing tasks (retrofits, not rebuilds)

Most RTOs do not need a rebuild; they need a targeted change to a task they already have. The commonest retrofits:

Existing task Typical weakness Low-cost alteration
Take-home knowledge questions Generic, fully outsourceable Add disclosure-and-annotation, or a short oral spot-check on a random sample of questions
Scenario-provided template The facts are supplied in the prompt Contextualise to the student’s own data, or require a specific detail only they hold
Written “how I would respond” for a performance skill Wrong evidence type Convert to an observed or recorded response with a live, responsive element
AI-banned first draft of a now-AI-assisted document Assesses a pre-AI workflow (obsolete) Convert to AI-draft plus annotate, correct and justify, then re-test authenticity
Scripted video presentation Tests delivery, not competence Add an unseen question or challenge the student must handle live

6.5 A rough cost and effort scale

Cost here means assessor time plus setup, not licence fees. The design principle is “cheapest that works,” so an RTO should reach for the lowest tier that genuinely closes the gap.

Effort Patterns Note
Low Contextualise-to-own-data; disclosure-and-annotation; process/iteration capture; compare-and-critique (async) Mostly wording and submission changes; little new assessor time
Medium Async recorded oral or video responses; disclose-and-defend; verify-and-source with a spot check; AI roleplay simulation Setup or batched grading time, but scales online and avoids a live-classroom cost
High Live synchronous observation or viva at scale; full workplace third-party sign-off coordination Strongest authenticity, highest coordination cost; reserve for where nothing cheaper suffices

The whole point of “verify once, strategically” is to spend high-effort verification only where the competency demands it, and to keep everything else in the low and medium tiers. In practice the binding constraint is usually assessor availability rather than dollar cost, so “cheapest that works” also means “least assessor-hours that work.”

6.6 Fairness and flexibility: the verification has to survive them too

Verification changes lean on live or observed components, and those engage the other two Principles of Assessment, fairness and flexibility, as directly as they serve the Rules of Evidence. A live oral disadvantages an anxious student, a learner for whom English is a second language, or a shift worker in another time zone; a synchronous video checkpoint assumes reliable connectivity a remote learner may not have. None of this is a reason to abandon verification, but it is a reason to prefer the cheap, flexible forms of it: asynchronous recorded responses over fixed-time live orals; a short spot-check on a sample rather than a universal viva; workplace third-party sign-off for work-connected learners; and, in every case, reasonable adjustment and unlimited low-stakes practice as standard. A verification design that is defensible on authenticity but fails fairness is not compliant either. Usefully, the “cheapest that works” principle and the fairness principle push in the same direction, toward the lightest sufficient anchor rather than the most secure one.

6.7 Two caveats on the require-AI fix

The judgement-over-output fix (section 5.5) carries two obligations of its own. Access and equity: if a task requires an AI tool, the RTO must name an acceptable free-tier option or provide access, so the assessment does not quietly test who can afford a paid model. Data handling: a task that captures the student’s prompts is putting student work into a commercial AI system, so the same privacy guidance that applies to the workplaces in these units (OAIC, 2024) applies to the method itself. Disclose the tool, keep personal or identifying information out of prompts, and check the tool’s data-retention terms. The recommendation to require AI and capture the prompt is not exempt from the obligations it teaches.

7. Limitations

This study is a method demonstration on a bounded corpus, not a survey of sector prevalence, and it should be read that way.

The corpus is a single provider’s productised builds off one template (the author’s own builds; see section 4.3). All seven packages were constructed to a common assessment template, which is why the AI-use policy is identical across them (section 5.1). This means the findings show what a systematic per-task audit detects, and what faults a shared template can carry, but they cannot establish how common these faults are across other RTOs’ independently produced materials. The prevalence question is untested here and would require a randomised sample of third-party assessment tools, which was outside the scope of this study.

The Resist/Integrate split rests on secondary sources, not primary workplace observation. Warrant establishes how AI is reshaping each occupational task from published regulator, national-body and industry sources (listed per unit in Appendix A), used as the best available proxy. The stronger evidence base would be direct observation of how AI is actually used in the relevant workplaces. That is impractical to coordinate at scale, dates quickly, and, importantly, sits outside an RTO’s remit: labour-market and workforce analysis is the role of bodies such as Jobs and Skills Australia and the Jobs and Skills Councils (Jobs and Skills Australia, 2025), not individual training providers. The method is designed to be refreshed as those bodies publish primary data, and its per-element calls should be treated as current-best-evidence, not settled fact.

The findings are a point-in-time snapshot (mid-2026). Both AI capability and the regulatory position are moving quickly. ASQA’s draft AI Principles were, at the time of writing, unpublished in full (section 3). A re-run in twelve months could reasonably reach different calls on the same units, which is a feature of the method, not a flaw: currency is one of its two axes.

The audit is AI-assisted expert judgement against a defined method, single-assessor. The Resist/Integrate calls and the per-task rulings were drafted by the AI tools described in section 4.4 and reviewed and owned by one qualified practitioner against a documented rule set; they were not validated for inter-rater reliability across multiple independent assessors. Full replication would require the calibrated thresholds behind the summary rules given in section 4.2, which are not reproduced in full; a lighter but real check is available, since a second assessor can re-run the two axes on the same units from the public unit requirements and the summary rules, and compare calls. The strongest next step, and the one that would show whether these fault patterns generalise beyond this corpus, is a larger study across many packages and multiple publishers with several independent assessors, measuring the inter-rater reliability of the Resist/Integrate calls and the rulings rather than asserting it. The method, not the seven-unit result, is what this paper offers others to test.

The sample is small and purposive (seven units, three faculties). It was chosen to demonstrate range, not to be statistically representative. Trade, health, and hands-on practical units in particular are under-represented, and their observed-performance assessment conditions may behave differently under this method.

One legitimate generalisation does follow, and it is the reason this matters beyond one provider. These packages were productised, template-built commercial resources of exactly the kind many small RTOs buy rather than build. If template-built commercial materials can carry these faults, then purchased assessment tools are precisely what the Standard 1.3 pre-use review exists to catch, and an RTO that has licensed resources from a publisher still owns the obligation to review and contextualise them before use. This does not establish prevalence, and it does not single out purchased materials as worse than in-house ones: the Standard 1.3 pre-use review applies equally to both. What this corpus shows specifically is that the purchased-template category can fail that review, so tools licensed from a publisher warrant the same scrutiny as anything built in-house, not the benefit of the doubt because someone else made them.

8. Recommendations

These are options, not a single mandated fix. Every RTO’s named subject-matter expert makes the final call for their own cohort and delivery mode. Each recommendation follows from the findings, and every one sits inside Standard 1.4.

  1. Make the AI-use statement operational, not decorative. Replace the blanket permission-and-prohibition most packages already carry with a two-tier policy that requires AI, with the prompt logged, on the tasks where it belongs, and prohibits it, with disclosure, everywhere else, so the policy and the task design finally agree.

  2. Audit at the element level, not the unit level. Several units were genuinely mixed, some elements to resist and others to integrate, inside the same unit and sometimes the same task. A single “our unit is or is not AI-proof” judgement misses this every time.

  3. Design one minimum-sufficient live or observed checkpoint per package, if none exists. This was the difference between a package that held up and one that did not. It need not be elaborate: a short oral defence, a recorded conversation, a supervised spot-check on a sample of questions. Zero verification anywhere was the most common fault found and the one with the clearest, cheapest remedy.

  4. Where AI is genuinely part of the modern job, assess judgement over the output, not the output itself. Require the tool, capture the prompt, mark on the correction and the justification, then re-test authenticity so the modernisation does not reopen an outsourcing hole.

  5. Do not manufacture AI-integration where the work does not call for it. The majority of elements in this corpus were left unmodernised. A response driven by anxiety rather than analysis risks bolting AI onto units where it has no legitimate place, which weakens validity rather than protecting it.

  6. Distinguish a systemic fault from a cluster of unrelated faults before choosing a fix. Some packages need one redesign decision; others need several different, task-specific patches. Treating the first as the second wastes effort; treating the second as the first hides genuinely different problems under one umbrella.

  7. Treat this as a Standard 1.4 exercise, not an AI-policy exercise. RTOs already validate assessment against the Rules of Evidence. This is that same obligation applied task by task instead of unit by unit, and it can be run now, under a Standard you already have to meet, without waiting for ASQA’s draft AI Principles to be finalised.

These recommendations are not a new process to run alongside existing obligations. The per-task audit is, in practice, the Standard 1.3 pre-use review for a new or purchased tool, and re-running the research stage on a schedule is part of the validation cycle the Standards already require for tools in use. The method is meant to run inside the machinery an RTO already operates, not beside it.

9. Conclusion

The useful way to treat AI and assessment is not as a new problem awaiting new rules, but as an old obligation under new pressure. Authenticity and currency have been in the Rules of Evidence all along. What generative AI has done is make it cheap to produce assessment evidence that satisfies neither while appearing to satisfy both, and, for the first time, cheap to check a whole package for it. This audit found a small set of faults recurring across seven packages built to one template: acceptable-use statements doing no real work, missing verification, scenarios handed to students, and written responses standing in for demonstrated skills. Whether those faults are as common in independently built materials elsewhere is the open question this study cannot answer and does not claim to. What it can show is that the faults are detectable, the fixes are specific and mostly cheap, and none of them requires the regulator to act first. They require reading the existing Standard as though it meant what it says.


References

Web sources accessed July 2026.


Appendix A: the seven worked audits

Each unit was researched at the element level (the Warrant Resist/Integrate split) and then ruled task by task on the two axes (Ruling), AI-assisted and human-reviewed as described in section 4.4. The tables summarise the calls and the reasoning. The summary rules a second assessor would apply are given in section 4.2; the calibrated thresholds behind them are the method’s detailed layer and are not reproduced here. Each unit lists its training.gov.au record and the primary sources behind its split.

A.1 BSBWHS411: Implement and monitor WHS policies, procedures and programs (the mixed unit)

Split. Resist: legislative accuracy, team consultation, hierarchy-of-control judgement. Integrate: drafting WHS communications, analysing inspection and incident data, hazard identification from digital tools, aggregate WHS data.

Task Axes Decision Reason
Facilitate a WHS team consultation Authenticity high (if observed), current Compliant Live consultation is AI-resistant and the split says resist. Confirm it is observed, not written.
Plan and provide WHS information Authenticity fine, currency obsolete Revise (modernise) Drafting WHS comms is now AI-augmented; assess judgement over an AI-drafted brief.
Identify and plan WHS training needs Authenticity fine, currency obsolete Revise (modernise) AI-assisted analysis of training-needs data is the real task.
Identify and report hazards Authenticity fine, currency obsolete Revise (modernise) Digital inspection tools are standard; assess the AI hazard sweep, then the correction.
Assess risk, apply hierarchy of control Authenticity medium, current Revise (authenticity) The judgement must be the student’s; add an unseen oral or recorded verification. Currency correctly silent.
Use aggregate WHS data Authenticity fine, currency obsolete Revise (modernise) Record platforms now analyse with AI; assess the override and justification.
AI-use policy (package) n/a Revise (policy) Blanket ban contradicts the modernised tasks; apply the two-tier policy.

Teaching point. The clearest demonstration of both axes firing in one package: currency drives the integrate tasks, authenticity does the work on the resist tasks, and modernising the tasks forces a policy rewrite.

Sources. training.gov.au/training/details/BSBWHS411; Safe Work Australia, Australian Work Health and Safety Strategy 2023-2033; SafeWork NSW, Ethical use of artificial intelligence in the workplace (final report); Safe Work Australia, AI and automated decision-making at work; TEQSA (2025).

A.2 CHCDIV001: Work with diverse people (the AI-thin unit)

Split. Every element Resist. Self-reflection on one’s own bias and real-time culturally responsive communication are not AI-augmentable on the student’s behalf.

Task Axes Decision Reason
Structured self-reflection on own perspectives Authenticity low, current Compliant Genuine reflection on one’s own biases is personal and specific; AI-thin, so unaided is current.
Respond to three real situations (verbal and non-verbal) Authenticity medium, current Rebuild Performance Evidence requires verbal and non-verbal communication; a written “how I would respond” is the wrong evidence type. Move to observed or recorded response.
Plan language support and escalation Authenticity low, current Compliant Anchored to specific situations and the group’s own policy.
AI-use policy n/a Compliant “AI to support learning, not as a deliverable” is consistent here; no task requires AI.

Teaching point. The guard against the false positive. The audit did not invent an “add AI” fix for a unit where AI does not touch the work. The only real fault was an evidence-type problem, caught on the authenticity axis, with currency correctly silent.

Sources. training.gov.au/training/details/CHCDIV001; Jobs and Skills Australia, Generative AI to augment and advance the way we work; HumanAbility, Submission to the Generative AI Capacity Study; Australian Human Rights Commission, AI and recruitment compliance checklist; Charles Sturt University, Cultural considerations in AI use.

A.3 BSBCRT411: Apply critical thinking to work practices (the pile-up, partly rescued)

Split. Resist: the role of critical thinking, the core analysis and reasoning, the reflection. Integrate: information gathering, drafting and idea-surfacing for the proposal.

Result. 17 of 19 tasks independently failed the authenticity test for the same reason: generic or scenario-provided prompts, submitted as text, with nothing verifying the thinking is the student’s own. Rather than issue seventeen identical fixes, the audit named the systemic pattern once as a package-level flag for comprehensive review, with a delivery-mode assumption noted (online and unmonitored, the harder case). Two things were carved out and called separately: the one recorded stakeholder presentation is genuinely Compliant (live, unscripted, responsive), and it anchors the written proposal draft as triangulated evidence (the product plus a live demonstration of command over it), enough that the draft needs only a Revise (modernise) for currency. The AI-use policy is a standalone Revise.

Teaching point. Even a mostly-broken package usually has a clean spot. Finding the one live checkpoint that already works, and using it to compensate the tasks around it, is as important as naming the systemic fault.

Sources. training.gov.au/training/details/BSBCRT411; Jobs and Skills Australia, Our Gen AI Transition; Future Skills Organisation, Top 10 Insights 2025; Lee et al. (CHI 2025), The Impact of Generative AI on Critical Thinking (international, flagged); TEQSA (2025).

A.4 BSBPEF201: Support personal wellbeing in the workplace (RESIST-heavy, real fault not about AI)

Split. Resist: wellbeing self-reflection, communication planning, the live conversation and its review. Split: the resources element, where AI wellbeing apps have joined the pool of resources a learner might find.

Result. The knowledge questions all failed the authenticity test on generic third-party framing with no verification anchor, collapsed into a single flag for comprehensive review rather than five separate fixes. The project’s two written steps read as Compliant through compensation: they feed a live, assessor-graded role-play that anchors them as triangulated evidence (the written plan plus a live demonstration of command over it), which is what the competency requires. The role-play itself is Compliant. The resources record gets a Revise with a two-part fix: a specificity anchor for authenticity, and an AI-app evaluation lens (not the standard judgement-over-output fix) for the narrow split. The AI-use policy is a Revise.

Teaching point. The headline fault here has almost nothing to do with AI. It is a plain authenticity design weakness (nameless scenarios, no anchor) that AI merely makes cheap to exploit. Not every problem the audit surfaces is an AI problem.

Sources. training.gov.au/training/details/BSBPEF201; Black Dog Institute, Using AI chatbots for mental health; Beyond Blue, Is AI safe for mental health support?; Safe Work Australia, Model Code of Practice: Managing psychosocial hazards at work.

A.5 BSBWRT411: Write complex documents (the total wash)

Split. Resist: audience and purpose, source verification, review and sign-off. Integrate: outline and structure, drafting the text, the mechanical proof and format pass.

Result. The as-built package engaged with none of this. It was built AI-free, take-home and artefact-only, with zero verification anywhere across all three assessments, on a unit whose most AI-exposed cluster (drafting) it assessed as a banned, unaided first draft. Every one of the six task clusters fired a fix-shaped fault. Because the package has no secure point at all, patching six tasks in isolation would still leave nothing anchoring the evidence, so the audit resolved to a single package-level flag for comprehensive review around one centrally designed verification point.

Teaching point. The case that proves naming a systemic fault has teeth. Six defensible individual fixes do not add up to a defensible package when the underlying fault is one design decision (build the whole suite unverified) repeated six times.

Sources. training.gov.au/training/details/BSBWRT411; Jobs and Skills Australia, Our Gen AI Transition; Fair Work Commission draft guidance on generative AI use (reported via Human Resources Director); CSIRO / National AI Centre, Guidance for responsible AI adoption; TEQSA (2025).

A.6 CHCLEG001: Work legally and ethically (the split unit)

Split. Resist: the ethical-judgement core and the breach-recognition half of the legal element. Integrate: finding and interpreting legislation, and workplace-improvement recommendations.

Result. Sixteen of seventeen knowledge questions and the project’s first three steps failed on authenticity (fully take-home, facts supplied, no verification), collapsed into one systemic Revise rather than many. The three situation-response steps additionally hit Rebuild, because a written “how I would respond” template is the wrong evidence type for the observed problem-solving the Assessment Conditions require. Only two spots fired the currency axis, and exactly where the split predicted: the legislation-lookup question and the practice-improvement step both got a Revise (modernise) with the judgement-over-output fix. The AI-use policy is a Revise.

Teaching point. Restraint and modernisation inside the same package. The audit modernised only the two integrate-flagged spots and left the resist two-thirds on authenticity grounds, while catching that “resist” governs what stays unaided, not whether an unaided description can stand in for a demonstrated performance. Those are different questions, answered on different axes.

Sources. training.gov.au/training/details/CHCLEG001; HumanAbility, Submission to the Generative AI Capacity Study; Department of Industry, Science and Resources, Australia’s AI Ethics Principles; Australian Association of Social Workers, AASW Code of Ethics 2020.

A.7 ICTICT451: Comply with IP, ethics and privacy policies in ICT environments (integrate-heavy, cluster-specific)

Split. Integrate: locating and reading organisational policy, drafting recommendations, documenting incidents. Split: non-compliance incident identification (documentation integrates, the genuine-breach call resists). Resist: the underlying lawful and ethical judgement.

Result. Nearly every task needed a Revise: the integrate clusters got the judgement-over-output fix plus a re-attached verification anchor (because each is also a scenario-provided take-home, so modernising alone leaves the authenticity gap open), and the resist question got an authenticity-only anchor. The audit deliberately did not invoke a systemic flag despite near-universal revision, because the fixes are genuinely different task to task rather than one repeated fault. The one uniform fault, zero verification across all fifteen items, was named on its own terms as a flag for the human resourcing decision of where to add a live component. A deliberate check for the classic IT-unit faults (written roleplay, knowledge substituted for performance) found neither: this is a compliance and judgement role with genuinely documentary evidence, so written templates are the right evidence type.

Teaching point. Two lessons. First, near-universal revision is not the same as a systemic fault; the fixes here are cluster-specific, and naming one blanket pattern would hide that. Second, a technical unit does not automatically require performance-style evidence: the archetype only bites when the unit demands demonstrable hands-on work, which this one does not.

Sources. training.gov.au/training/details/ICTICT451; Office of the Australian Information Commissioner, Guidance on privacy and the use of commercially available AI products; Department of Industry, Science and Resources, Australia’s AI Ethics Principles; Allens, Using gen AI tools for business: limitations on your IP protections; Jobs and Skills Australia, Our Gen AI Transition.