Gen AI Lingo

Your next step

Test reliability beyond a headline score

Fair conditions, contamination, robustness, calibration, judges.

Start learning

What you’ll learn

  1. Keep evaluation conditions comparable

    Identify an unfair model comparison.

    About 4 minutes · Open activity

  2. A memorized exam is weak evidence

    Detect benchmark contamination.

    About 4 minutes · Open activity

  3. Robustness means handling relevant changes

    Choose a meaningful perturbation for a test.

    About 4 minutes · Open activity

  4. Confidence needs calibration evidence

    Distinguish an evaluated confidence estimate from confident wording.

    About 4 minutes · Open activity

  5. An explanation may be unfaithful

    Separate a useful answer explanation from a verified account of computation.

    About 4 minutes · Open activity

  6. AI judges also need evaluation

    Identify a bias or criterion mismatch in an automated evaluator.

    About 4 minutes · Open activity

  7. Probe boundaries safely

    Recognize the purpose of adversarial testing on fictional material.

    About 4 minutes · Open activity

  8. Audit a model demo

    Find an unfair test condition, a weak evidence claim and an untested reliability boundary in a fictional product demonstration.

    About 5 minutes · Open activity