AI Fluency Assessment: How to Test Your Employees' AI Skills
Twenty scenario questions across the four dimensions of AI fluency, scored to a proficiency level per dimension. Take it in Vero in one sitting, or run the questions yourself.
An AI fluency assessment measures how well an individual employee delegates work to AI, describes what they need, judges what comes back and uses the tools responsibly. It scores judgment on realistic scenarios, not tool knowledge or self-reported confidence, and reports a proficiency level for each of the four dimensions.
What an AI fluency assessment measures, and what it does not
It measures a person, not a company. Our AI readiness assessment scores your organization on strategy, governance, data, talent, adoption and measurement. This one scores an individual employee on what they would actually do with AI at their desk, in your tools, under your data rules. You need both, and they answer different questions.
It is not a hiring screen. Interview rubrics score a candidate's account of past work. This assessment puts a current employee in front of realistic situations and scores the judgment they show, so the result is evidence rather than narrative.
It measures four habits rather than tool knowledge. The questions follow the AI Fluency Framework developed by Rick Dakan and Joseph Feller, the same framework Anthropic teaches in its AI Fluency course. Its four dimensions are delegation (deciding what to hand to AI and what to keep), description (communicating the task, its context and its constraints), discernment (judging the output, the process and the tool's behavior critically) and diligence (using AI responsibly: data, disclosure and accountability).
The reason to assess habits instead of tools is durability. Copilot, ChatGPT, Gemini and Claude change every quarter. The four habits do not, so a baseline taken today is still comparable in a year, and a person who scores Advanced transfers that fluency to whatever tool your company standardizes on next.
The scoring rubric: four levels, four dimensions
Each dimension is scored to one of four proficiency levels: Beginner, Intermediate, Advanced or Expert. These are the same four levels Vero reports for every skill, so a fluency result sits alongside the rest of an employee's skill profile without translation. Score each dimension separately. A person is rarely at the same level on all four, and the spread is the most useful thing the assessment tells you.
| Dimension | Beginner | Intermediate | Advanced | Expert |
|---|---|---|---|---|
| Delegation | Uses AI when told to, or not at all. Cannot say which of their tasks suit it. | Hands AI the obvious tasks: drafts and summaries. Keeps or delegates by habit, not judgment. | Splits a task into AI-suited and human-only parts and can explain why for each. | Redesigns a workflow around the split and sets the delegation rules others follow. |
| Description | One-line requests. No context, format or constraints. | Adds context and a format. Rarely states constraints or success criteria. | Gives audience, context, source material, constraints, format and an example. Iterates on a weak output rather than starting over. | Builds reusable instructions and templates the team adopts. |
| Discernment | Accepts output as produced. Cannot name a way it could be wrong. | Spots obvious errors. Misses fabricated citations, stale facts and subtle bias. | Verifies claims against a source before use. Notices when the reasoning, not just the answer, is off. | Designs the checks (sampling, source rules, review gates) that other people run. |
| Diligence | Unsure what data is permitted. Would paste customer data into a consumer tool. | Knows the policy exists. Applies it inconsistently under time pressure. | Follows data rules unprompted, discloses AI use where it matters, owns the output as their own. | Anticipates new risks, escalates them and shapes the policy. |
The 20 questions
Five questions per dimension. Eight are multiple choice and twelve are open answer, and each carries a difficulty from Beginner to Expert. In Vero, the multiple-choice items present four options; the option sets are not reproduced here. Every question is shown below as a prompt, with a note on what a strong answer shows, because the thing being scored is a way of thinking rather than a keyword.
The material embedded in the questions (a task list, meeting notes, a paragraph with planted errors) is fixed, so results are comparable across people and teams. If you run the questions yourself, keep the material as written for the same reason.
Delegation — Deciding what to hand to AI and what to keep.
1. Here is your task list for tomorrow morning: (a) reply to a customer asking for an exception to the refund policy; (b) summarize three competitor press releases; (c) draft the first version of a project status update for your manager; (d) decide which of two vendors to recommend; (e) reformat a spreadsheet of contacts into a clean table. Mark each item AI-first, AI-assisted or human-only, and give a reason for each.
Open Answer · Intermediate
A strong answer shows: A reason per item, not a blanket rule. Advanced answers weigh risk, reversibility and whether a source exists to ground the work.
2. A colleague uses AI to write every customer email end to end. Which approach shows the best judgment about where to draw the line?
Multiple Choice · Beginner
A strong answer shows: Routine updates can be drafted with AI; sensitive or relationship-critical messages are written by the person. Neither “never” nor “always” is a strong answer.
3. You have 20 minutes to produce a competitor summary for a sales call. Describe, step by step, how you would split the work between yourself and an AI tool, and say what you would keep for yourself.
Open Answer · Advanced
A strong answer shows: The person keeps the framing, the judgment and the final call, and delegates collection and first drafts.
4. Name two tasks in your own role where using AI makes the result worse, and describe what the failure looks like in each case.
Open Answer · Intermediate
A strong answer shows: Specific, role-based examples. Beginners cannot name one. Experts name the ones their team gets wrong.
5. Your manager asks you to “use AI more.” Which is the best first move?
Multiple Choice · Intermediate
A strong answer shows: One recurring, low-risk task, tried and measured, rather than moving everything at once, waiting for formal training or running approved prompts verbatim.
Description — Communicating the task, its context and its constraints.
6. Write the exact request you would give an AI tool to produce a one-page brief for your leadership team on whether to move the weekly all-hands meeting to every two weeks. Then assume the first output is too generic, and write the follow-up message you would send to improve it without starting over.
Open Answer · Intermediate
A strong answer shows: Audience, context, format and constraints in the first attempt. A targeted follow-up rather than a fresh request.
7. The output is generic and misses your company's terminology. What is the most effective change to your request?
Multiple Choice · Beginner
A strong answer shows: Supplying a real example and the terms the team uses, rather than repeating the request with more emphasis, asking for “more specific,” or switching tools.
8. Which set of elements makes a request most likely to produce a usable result on the first attempt?
Multiple Choice · Beginner
A strong answer shows: Audience, context, source material, constraints and the output format needed.
9. Here are raw notes from a client call: “budget approved but only through Q1 – Dana wants weekly check-ins not monthly – integration w/ their CRM still blocked on their IT, ETA unknown – they liked the onboarding video, want a version for managers – pricing question on extra seats, said we'd come back by Friday – Dana's boss joining next call.” Write the instructions you would give an AI tool to turn these into a client-ready summary, and explain what you told it to keep, what to cut, and what the reader needs to decide.
Open Answer · Advanced
A strong answer shows: Instructions that state what must be kept, what must be cut and the decision the reader faces.
10. A teammate keeps getting worse results than you from the same AI tool. What are the three questions you would ask them, and why those three?
Open Answer · Intermediate
A strong answer shows: Questions about the context given, the constraints set and how they iterate. Experts describe a template they would hand over.
Discernment — Judging the output, the process and the tool's behavior critically.
11. The paragraph below was written by an AI tool. It contains two factual errors and one invented source. Identify all three and say how you would verify each. “Email remains the dominant channel for internal communication. The first email was sent in 1971 by Ray Tomlinson between two computers in different cities, and by 1995 more than half of US households had a home internet connection. A 2023 study by the Harvard Institute for Workplace Analytics found that employees spend an average of 28% of their working week on email. Tomlinson chose the @ symbol to separate the user's name from the host machine.”
Open Answer · Advanced
A strong answer shows: All three found, each with a verification method. Partial credit for two.
12. The tool gives you a confident answer that disagrees with what you remember. What do you do?
Multiple Choice · Beginner
A strong answer shows: Check a primary source before acting; the tool's confidence is treated as no evidence either way.
13. Which list names three ways AI output can be wrong that a spelling and grammar check would never catch?
Multiple Choice · Intermediate
A strong answer shows: Invented citations, outdated facts and reasoning that sounds plausible but fails, as opposed to surface errors of spelling, style or formatting.
14. An AI-drafted analysis reaches a conclusion you already agree with. Does it still need checking? What specifically would you check, and why?
Open Answer · Intermediate
A strong answer shows: Yes. The reasoning path and the numbers, not just the conclusion. Names confirmation bias or its equivalent.
15. You are reviewing a colleague's AI-assisted work before signing it off. Describe your review process step by step.
Open Answer · Expert
A strong answer shows: A repeatable method: sample the claims, verify sources, test edge cases. Experts describe a gate others could run.
Diligence — Using AI responsibly: data, disclosure and accountability.
16. Which is the safest general rule for deciding what company data may go into an AI tool?
Multiple Choice · Beginner
A strong answer shows: Follow the written policy for that tool and ask before pasting anything it does not cover. “Visible to everyone,” “delete the chat afterwards” and “paid tools are safe” are all wrong.
17. A customer's contract is in your inbox and a consumer AI tool would summarize it in seconds. Talk through your decision.
Open Answer · Intermediate
A strong answer shows: The data question is raised before the convenience. The person knows there is a sanctioned alternative.
18. You used AI heavily on a deliverable. Do you tell the recipient? When, and how?
Open Answer · Advanced
A strong answer shows: Disclosure calibrated to the stakes and the audience. Ownership of the result regardless.
19. An AI-assisted email you sent contained an error that reached a customer. Whose responsibility is it, and what happens next?
Multiple Choice · Intermediate
A strong answer shows: The sender's: correct it with the customer, then change the check before sending. Not the vendor's, not IT's, and not nobody's because AI was disclosed.
20. You notice a new AI feature in a tool the team uses that could expose company data. What do you do?
Open Answer · Expert
A strong answer shows: An escalation path, a pause on use and a willingness to shape the rule. This is the clearest Expert signal in the set.
How it is scored
Every question belongs to one dimension and carries a grading note that describes what a strong answer covers. Responses are scored against those per-question notes rather than a single answer key. Multiple-choice items count for correctness weighted by difficulty; open answers are judged on depth, accuracy and practical understanding.
The output is a proficiency level for each dimension (Beginner, Intermediate, Advanced or Expert), with two to four strengths and two to four areas for improvement under each, plus an overall level averaged across the four. The per-dimension levels are the result that matters. The overall level is a summary; the spread is the diagnosis.
Read the spread before the headline. The example below is typical: strong on description, weak on discernment. The training plan writes itself from the lowest dimension, not from the average.
Two reading rules worth adopting. First, train the lowest dimension first, whatever the overall level says, because fluency is a chain and an employee who describes tasks like an Expert but discerns like a Beginner produces polished, unverified work at speed. Second, treat a Beginner result on Diligence as the priority regardless of the other three: data exposure is the one failure with no upside, so policy and sanctioned tools come before anything else.
For leadership, the three numbers worth reporting are the share of people at each level per dimension, the movement between assessments and the application evidence (verification catch rate, quality of AI-assisted output). Completion rates and satisfaction scores are not fluency metrics.
How to run it
In Vero, the General AI Fluency Assessment is a catalog assessment in the Foundational track: 20 questions, about 40 minutes, completed in a single sitting with no time limit. Questions are presented one at a time with a progress bar. On submit, every response is graded in one pass and the employee sees an Assessment Results screen with their overall level, a level for each of the four dimensions, and strengths and areas for improvement under each. The result is written to their skill profile, so it shows up on the skills dashboard next to everything else they have been assessed on.
Managers see Workforce Results: who is assigned, in progress and complete, the overall level distribution, and each dimension broken out, so a team's weakest dimension is visible in one view.
- Running it without Vero. Give the 20 questions as a written exercise, about 40 minutes, with the embedded material kept as written. Score each answer to a level using the rubric row for its dimension, take the typical level across the five questions, and record four levels per person. Have two people score the first ten sessions independently and compare, so the rubric means the same thing across teams.
- Self-assessment. Use it only as a pre-read. Self-rated confidence runs high and measures confidence, not capability. If the self-rating and the scored level disagree, the scored level is the one you record.
- Sample. Assess the whole team if it is under 50 people. Above that, start with a 20-person sample per function to get the distribution, then extend.
- Cadence. Baseline before any training. Reassess at 90 days, then quarterly for AI-central roles and twice a year for everyone else. Keep the questions stable so movement is real.
- Share the levels, not the answers. People should know what Advanced looks like before they sit down. The questions themselves being public is fine: they test judgment, not recall, and knowing you will be asked where to draw a line does not tell you where it is.
What to do with the results
Each level has a different next step, and getting the order wrong wastes the budget.
- Beginner: policy and sanctioned tools first, then guided practice on two real tasks from their role. Prompt tips at this stage teach people to produce faster work they cannot check.
- Intermediate: role-specific practice with feedback on description and discernment. This is where the largest jump usually happens, and where interactive practice beats a video library by the widest margin.
- Advanced: stretch tasks, workflow redesign and peer teaching. These people are your scorers for the next round.
- Expert: put them on the assessment panel and the policy group. Their job is to raise the levels around them.
If you want a schedule to hang this on, our 30-60-90 day AI fluency program template starts with this baseline, and our guide to measuring AI fluency covers the reporting.
How OneRange Vero runs this at scale
Scoring twenty open answers by hand works for 30 people. At 300, scorers drift, the sessions eat a quarter and the results stop being comparable across teams. In OneRange Vero the General AI Fluency Assessment is a ready-made skill assessment from the Vero AI Catalog. Each question is mapped to a skill in Vero's 10,000+ skill taxonomy, one per dimension (Identifying AI Use Cases, Prompt Crafting, AI Output Interpretation and AI Policy Compliance), so every employee gets a proficiency level on each dimension in one sitting and the level means the same thing in sales as it does in finance.
Then Vero generates interactive AI training from your own authoritative source material, with citations. Your data policy, your tools, your workflows. Training is assigned against the levels the assessment found, so a Beginner gets the policy and guided practice while an Advanced employee gets stretch work, and reassignment after training measures movement rather than assuming it. Unlimited adaptive AI sessions. Launch in under two weeks.
See the assessment in the Vero AI Catalog
Frequently asked questions
What is an AI fluency assessment?
A structured test of how an individual employee delegates work to AI, describes what they need, judges the output and uses the tools responsibly. It scores judgment on realistic scenarios and reports a proficiency level for each of the four dimensions, with strengths and areas for improvement under each.
Can you assess AI fluency with a survey?
Not usefully. Surveys measure confidence, and self-rated confidence runs high and does not predict capability. Use a survey as a pre-read if you like, but record the level from scored answers.
How long does an AI fluency assessment take?
About 40 minutes for the 20 questions, completed in a single sitting. In Vero there is no time limit and results are available as soon as the responses are graded.
What is the difference between AI fluency and AI literacy?
Literacy is understanding what AI is and how it works. Fluency is using it well inside your own job, with judgment about when not to, and verifying what it produces. This assessment tests fluency; you can be literate and still score Beginner.
Do we have to use the four Ds?
No, but they are a good structure because they describe habits rather than tools, so the baseline stays comparable as tools change. They are the same four dimensions Anthropic teaches in its AI Fluency course, which makes them familiar to many employees already.
Should employees see the rubric before the assessment?
Yes. Share the four levels and what each looks like so people know the standard. The questions test judgment rather than recall, so seeing them in advance does not change what a strong answer requires.
Keep reading
- General AI Fluency Assessment in the Vero AI Catalog — the assessment itself, ASM-FLU-07
- What is AI fluency? — the definition, the four dimensions and how to design the program
- How to measure AI fluency — baselines, levels and the three numbers to report
- 30-60-90 day AI fluency program — a schedule that starts with this baseline
- AI readiness assessment for companies — the organization-level score, in 24 questions
- Skill assessments in OneRange — how Vero scores proficiency across 10,000+ skills
Run the baseline. Then train the gap.
OneRange Vero gives every employee a proficiency level on every dimension in one sitting, then builds interactive AI training from your own source material to move them up.