Science & validation

Science before scores.

genAssess was built to answer a hard question: can we measure the knowledge people need to work with AI in a way that is reliable, explainable and useful?

Here is the evidence so far, including its limits.
900+research hours
180initial items
289pilot participants
152refined items

Development history

From expert model to operational assessment.

The work began in mid-2024, combining talent assessment, responsible AI, governance and applied workplace expertise.

01

Define the construct

More than 900 hours of research, practitioner interviews and standards mapping shaped an expert model of the knowledge behind AI readiness.

02

Build and refine the bank

An initial 180-item bank was piloted and reviewed for clarity, difficulty and distractor performance, leaving 152 refined items.

03

Test the evidence

Pilot data from 289 participants was analysed for internal reliability, construct alignment and differential item functioning.

04

Operationalise fairly

Balanced testlets and common anchor items support varied but comparable 48-question assessment forms delivered through Sova.

Psychometric evidence

Four questions every assessment should answer.

Evidence is presented plainly so buyers, psychologists and governance teams can decide what weight to place on a score.

Reliability

Does it measure consistently?

KR-20: 0.85–0.91

Pilot scales reported strong internal consistency, indicating that items within the measure are working together.

Convergent validity

Does it align with an independent measure?

r = 0.83 with GLAT

Scores showed a strong relationship with the academic Generative AI Literacy Assessment Test.

Fairness

Do items behave consistently across groups?

DIF analysis completed

Items were examined for differential functioning across age, gender and ethnicity within the pilot sample.

Decision design

Is the score used proportionately?

Human decision retained

The assessment is designed as one input alongside other evidence. It does not make hiring decisions autonomously.

What the structure tells us

One overall construct. Six useful lenses.

Early factor analysis indicates a strong general AI knowledge factor. The six Core-6 scores remain valuable diagnostic lenses for interviews, learning and workforce planning.

General AI
knowledge
AIFAUCBERDQGOIEHOC

Responsible interpretation

Strong early evidence is not the end of validation.

What is established

The pilot provides evidence of internal reliability, convergent validity and considered item-level fairness.

What continues

Criterion validation with launch clients will examine how scores relate to meaningful workplace outcomes over time.

Put the model to work

See the assessment in context.