Back to News
July 25, 2026

Phishy or Not? vs the whole "phishing test + module" model

A different approach to human risk

Written by Lively, the team behind Phishy or Not?

TL;DR

  • What it is. Phishy or Not? replaces the whole "phishing test + module" model. Instead of catching who clicks and marching them off to a module, it builds and measures scam-spotting instinct through play.
  • The problem with the model. Simulate an attack, catch the click, send the failer to a module. MITRE found it did nothing (2014). ETH Zurich found it made some people more susceptible (2022). A 19,500-person UC San Diego study found a 2% effect, with most training closed within 10 seconds (2025).
  • The threat has outrun the model. AI has stripped away the old red flags, and attacks now arrive by SMS, voice, QR and deepfake, channels an email test never touches. A model that was marginal against yesterday's phishing measures is even less of today's.
  • The problem with the metric. The vendor writes the test, sets its difficulty and is judged on your click rate. Easy tests produce falling numbers; falling numbers renew contracts. NIST built the Phish Scale because click rates alone mislead.
  • The approach. Short scenario-based play across email, SMS, voice, QR and deepfakes, on any phone, for desk and non-desk staff alike. It measures behaviour rather than clicks.
  • The proof. 83% more reporting, 2x fewer clicks against benchmark, 9.8/10, and 95% choose to play on.
  • Compliance. No framework mandates simulation. They require tested awareness, which this meets, with per-person evidence. DORA red-team testing is the one exception.

That's the what. The why is the good bit, and it starts in 2004, with the US Army, a cannon, and a test more than 80% of people failed. Read on.

(Or ask your favourite LLM to summarise it. We wrote it to work either way.)

People keep asking what makes Phishy different

Is it a new category, or an extension of what's already out there? And who's it actually for?

Fair questions. To answer them, you need a bit of history.

The first well-documented "friendly phish", sending your own people a fake scam to see who bites, wasn't a software vendor. It was the US Army. In 2004, West Point ran an experiment called Carronade: a fake email from a made-up colonel, sent to 512 cadets, to see who'd click. More than 80% did. Among first-years who'd just sat through four hours of computer security instruction, it was 90%.

But here's the part everyone forgets. Carronade was named after a short-range cannon on purpose. The point was to help, not humiliate. A wake-up call. You clicked, you learned. No shame, no punishment.

Somewhere along the way, the model changed. The "friendly phish" became a covert gotcha: trick your staff, log who failed, send them off to a module. Most of the big vendors still run it that way.

Our co-founder, Stacey Edmonds, spotted in 2015 that it wasn't working, and built something else. Funny thing: it landed closer to what the Army actually meant. Wake people up. Build the instinct. No shame.

So, to answer the question, Phishy or Not? extends what Carronade intended and replaces what the industry turned it into. And the audience? Everyone.

Beyond compliance

Phishing simulation and security awareness training, the method behind the incumbents, was built to test employees at a desk. It measures who clicks.

Phishy or Not? takes the same job beyond the desk: scenario-based play that builds cyber safety instinct in everyone, everywhere, every channel, at work and at home, no IT project.

The metric grades its own homework

Here's the mechanism that keeps the model alive.

A click rate is not a property of your workforce. It's a property of the test. NIST built the Phish Scale in 2020 for exactly this reason: click rates rise and fall with the difficulty of the lure, so the same workforce produces a low click rate on an easy phish and a high one on a hard phish. Analysed on their own, NIST warns, low click rates can create a false sense of security.

Now look at who controls the difficulty. The vendor writes the lures. The vendor schedules the campaigns. The vendor publishes the benchmark you're compared against, generated from tests on its own platform. And the vendor's renewal depends on your click rate going down. Nobody in that loop is rewarded for sending a harder test. A falling click rate can mean your people got sharper, or that the tests got softer, and the report doesn't tell you which.

The outcome research says which. MITRE ran the experiment in 2014: no significant effect, and most employees never read the training. ETH Zurich ran it for 15 months across 14,000 employees: embedded training didn't help and made some people more susceptible; the authors cautioned against deploying it. UC San Diego ran it across 19,500 people in 2025: a 2% effect, with over half of training sessions closed within 10 seconds. And a 2024 review of 42 studies found the things that do improve detection are training intensity, active learning and detailed, immediate feedback. Not one of those describes a click-and-module loop.

Two more nails. Whatever a module does teach is gone within about six months, which means an annual cycle leaves your workforce uncovered for half of every year. And a 2026 analysis found the model can't even measure itself: click-triggered training is only ever assigned to people who already clicked, biasing every effectiveness claim built on it, and consequence-free tests showed an emboldening effect, teaching persistent clickers that clicking has no cost.

If you're running this model today, none of this is a verdict on you. It was a rational buy when work lived in an inbox. The threat moved to SMS, voice, QR and deepfakes, and the model and its metric never followed.

But that's what it is now, not a measurement system. Compliance theatre: a number that performs risk reduction for the audit without producing it.

Phishy or Not? exists for the other option: proof people changed rather than click logs, reach for the staff who never sit at an inbox, one platform across email, SMS, voice, QR and deepfakes, and human risk reporting your board and insurer can actually use.

Side by side

The incumbent model
The “phishing test + module” approach
The model shared across the security awareness category
The new model
Phishy or Not?
by Lively
Built forWork: a compliance task done at a desk.Life: a 24/7 instinct that goes home with them.
AudienceBusiness: employees with an inbox.Everyone, everywhere: every workforce, on their own phone, at work and at home.
Core modelSimulate an attack, catch the click, assign a module.Scenario-based play that builds instinct.
What it measuresExposure: click rate and completion.Capability: behaviour against The Signs, our behavioural taxonomy. Confidence vs competence.
Success metricA falling click rate on tests the vendor designs. Easy tests look like progress.Capability growth, with scenario difficulty on the record. The score can't be flattered by softening the test.
ReportingPer person, but it's a failure log: who clicked, who to retrain, who goes on the repeat-offender list.Per-person capability and calibration evidence, tracked over time. Board and insurer ready.
When you learnOnly when you fail: click, get caught, get a module.Every time: instant feedback on every decision, and you learn to spot what's safe, and why. No gotcha, no shame list.
ChannelsEmail core; some now add SMS, voice or QR.Native across email, SMS, voice, QR (quishing), deepfakes and real-world scenarios.
Threat contentScheduled library.Responds to threats daily; a new security alert can be a live in-game scenario the same day.
Personal benefitWorkplace-scoped. A one-off test, done and forgotten.A genuine employee benefit: instinct that protects their family and their own money too.
Data and privacyNeeds your staff directory and personal data to run.No directory sync or identity-system integration. In schools, the school holds the logins, so no student PII.
SetupAn IT project: mail-gateway allowlisting, directory sync, SSO.No email or identity-system integration required. SSO optional, live in days.
The experienceOff-the-shelf vendor content: the same templates every other customer gets.Customised to your organisation and crowd-sourced from the people who see scams first. White-label available.
Commercial modelCommercial software, priced per seat, sold to satisfy the audit.Social enterprise: every Phishy or Not? subscription helps fund Dodgy or Not? for schools.
Incumbent-model claims reflect vendors’ public statements about the shared approach, not any single product. Multi-channel is converging across the category; some vendors now add SMS, voice or QR.

The model isn't fading; it's automating. In February 2026 the category's largest vendor launched an autonomous agent that creates, schedules and assigns tests and modules on its own. A process that doesn't work has now been automated.

Where Phishy or Not? is genuinely different

It measures behaviour. A click rate tells you who fell for one email on one day, at a difficulty someone else chose. Phishy or Not? measures competence signal by signal, and whether each person's confidence matches their actual performance, the gap that predicts real-world risk. Scenario difficulty is coded against The Signs and sits on the record, so the score can't be flattered by softening the test. That's evidence your auditor, board and insurer can actually use.

This gap is documented, not a theory. In one 600-person study, 80% of people were more confident than their accuracy justified, and confidence barely predicted who could actually spot a phish. Decision scientists' recommended fix is measured feedback on actual ability, person by person. That's what the platform does.

It reaches the whole workforce. Email simulation only tests the people in an inbox. The nurse between rounds, the driver in the cab, the carer on shift, the member-services team on the phones. They're all targeted by text, voice and deepfake and never touched by an email test. Phishy or Not? runs on the phone in their pocket.

It teaches instinct that transfers. Every scenario is coded against The Signs, our proprietary behavioural taxonomy. Scam formats change constantly. The cues behind them, authority, urgency, loss, don't. Learn to read the signs once, and the instinct carries to every channel and every new format.

It stays current like an alert system. When a new scam circulates, it becomes a scenario the same day, drawn from government advisories and a network of CISOs who see attacks first.

It builds trust. No tricked employees. No walk of shame. The UK's NCSC warns that gotcha simulations erode trust for little security gain. People choose to play, and the instinct goes home with them.

It works because it's a game. Passive modules get skipped: in the UC San Diego study, more than half of training sessions were closed within 10 seconds, and fewer than a quarter of users finished the materials. The same study found the harm concentrated in static modules; interactive formats didn't show it. Engagement is the mechanism. Repeated exposure, instant feedback and active recall are how instinct forms and how it sticks. And 95% of players choose to finish the full game, unprompted. That's a habit-forming.

The compliance question, answered

The rules that apply to you, ISO 27001 (Annex A.6.3), PCI DSS v4.0 (12.6.3.1), NIST CSF 2.0 (PR.AT), HIPAA, NIS2 (Article 21) and APRA CPS 234, ask for the same thing: that your people are trained to spot and report social engineering and that you can prove it. Not one of them mandates phishing simulation by name. They require tested awareness. Scenario-based play meets that, and hands your auditor per-person behavioural evidence rather than a click log.

A falling click rate on vendor-designed tests is compliance theatre. Per-person capability evidence is compliance. It's also the stronger answer to the awareness-training line on your cyber-insurance questionnaire.

One  exception: DORA. Its threat-led penetration testing regime (Articles 26-27) requires designated financial entities to undergo red-team testing at least every three years, typically including live social engineering against staff. That's a red-team engagement no awareness platform satisfies on its own. If that's you, you'll want both, and we partner on it.

The evidence

Here's what our deployments show. People are 83% more likely to report a dodgy message, and reporting is the metric that matters: every report is an early warning your security team didn't have before. A 2x reduction in click rate, measured against the category's published benchmarks, not against tests we designed ourselves. A 9.8 out of 10 rating from the people playing it. And 95% choose to finish the full game, unprompted.

The numbers come from named deployments, including Transport for NSW, Toll and AICWA.

For education

Here for education rather than the workplace? Dodgy or Not? brings the same approach to students: curriculum-aligned, age-banded, and genuinely a game they ask to play again. It works the same way too: students show a 50% decrease in susceptibility to unwanted contact, measured in-game through behaviour, not a quiz. Every Phishy or Not? subscription helps fund it.

Our mission - to build cyber safety instinct in everyone, everywhere.

FAQ

Looking beyond KnowBe4? What are the options?

The usual list is Hoxhunt, SoSafe, Proofpoint, Cofense, Phriendly Phishing, Mimecast and Usecure. They differ in polish, price and geography, but they share the model: simulate, catch the click, assign a module, report the rate. If your problem is a vendor, switching solves it. If your problem is the model, switching swaps the logo on the same click report. Phishy or Not? is the option outside the model: it measures and changes behaviour across every channel.

How is Phishy or Not? different from phishing simulation platforms?

They catch the click and assign a module. Phishy or Not? builds instinct through scenario-based play, reaches non-desk staff on their own phones, and measures behaviour against a behavioural taxonomy rather than logging completions and click rates.

Does Phishy or Not? replace a phishing simulation programme for compliance?

For awareness, behaviour change and human-risk measurement, yes, and it maps to the common frameworks. It doesn't replace a DORA-mandated red-team penetration test, which is a separate engagement.

Do we have to take out our current phishing simulation platform?

No. Run Phishy or Not? alongside your existing program, put per-person capability evidence next to the click log and decide at renewal. We're comfortable with that comparison, and we can guarantee your people will be also.

Is it only for large enterprises?

No. It's built for all workforces, at their desk, on their feet, or at 20,000 feet: healthcare, aged care, trades, logistics and member organisations.

And it goes live in days without an IT project.

See it for yourself

For business: see the human risk report your board would actually get.

Book a Phishy or Not? demo

For education: Explore Dodgy or Not?

Sources