Outcome records
What we measured, and how
Aicoursesforsg publishes three classroom records. Each one is a timed exercise on anonymised sample data, taken in Fusionopolis in 2026, with a second person counting errors against a prepared key. The numbers describe that room, that sample, and that day. They are not a forecast of what your organisation will save.
Ask for a scoping call Course catalogue
The records exist so that a buyer can see the protocol before paying a fee. We publish a small, dated table with a denominator. If a figure on the homepage looks useful, this page is where you check the sample size and the exclusions.
How we measure
Every featured record uses the same skeleton. Learners complete the sample with their current method. Instructors then rebuild the procedure with a prompt file and a checklist. Learners complete a held-out slice of the same class of task. A second person, who did not write the prompt, scores the output against an answer key prepared before class.
Time is taken with a visible timer. The clock starts when the first item is open and stops when the learner declares the set complete. We record wall-clock minutes. Errors are counted as discrete events: a wrong label, a figure that does not match the source, a forecast cell that violates the stated method.
Two timed passes are the minimum. If a learner’s machine fails, that row is dropped. If a learner has never done the task in real work, that row is also dropped: the record is about a change in a known procedure, which requires a baseline that exists.
Three records, side by side
Handling time and error rates are shown as recorded. Sample size is the number of learners whose rows were kept after the exclusions above. Dates are the open-class or evening cohort in 2026.
| Task | Baseline | After | Sample | Date |
|---|---|---|---|---|
| Service inbox triage (20 tickets) | 6.8 min median; 9.0% mislabel | 2.9 min median; 3.5% mislabel | 12 | March 2026 |
| Monthly demand forecast (held-out quarter) | 18.4% MAPE | 11.2% MAPE | 9 | May 2026 |
| Board-pack reporting draft | 74 min median; 6.2 gaps | 31 min median; 1.8 gaps | 11 | June 2026 |
Open the full write-ups
Each case page repeats the protocol for that task, lists the files learners left with, and states what would have to be true to measure the same thing inside an organisation.
Service inbox triage
Median handling time on a 20-ticket anonymised sample. Mislabel rate moved from 9.0% to 3.5%.
18.4% → 11.2%Monthly demand forecast
Mean absolute percentage error on a held-out quarter of anonymised monthly series.
74 min → 31 minBoard-pack reporting draft
Median time to a complete first draft with source marks, plus second-marker gap counts.
Limits of these numbers
The samples are small. Twelve, nine and eleven people is enough to run a class and too few to support a general claim about “AI in Singapore offices”. The conditions are a training room with a prepared file, a visible timer, and an instructor in earshot. That is a different pressure from a live Monday with a queue and a manager waiting.
The data are anonymised demo packs. They are built to look like the work, with the same field names and the same messiness we see in scoping calls. They are still demo packs. A live inbox with attachments, a live forecast with a broken product hierarchy, or a live pack with last-minute director comments will move the stopwatch.
There is a novelty effect. People who have just been taught a method tend to follow it carefully for the second timing. We do not have a 90-day re-measure on these three cohorts, so we do not claim persistence. We also do not publish an average across all participants in all programmes: mixing Foundations drills with specialist tasks would hide the thing a buyer actually needs to read.
What we would need to measure it in your organisation
- A named task with a repeating sample. Twenty tickets of the same class, a year of a monthly series, or a pack that is produced on a fixed cycle. One-off projects cannot be re-timed.
- An anonymised extract you are allowed to use. If the file cannot leave the premises even after stripping, the measurement has to happen on-site with your own second marker, under a private-cohort contract.
- A second person who did not write the prompts. Scoring your own output is how errors get defined away. The classroom protocol uses a second marker for that reason.
- Agreement on what counts as an error before the clock starts. Labels, units, source marks, escalation cases. Changing the key after the after-pass invalidates the pair.
If those four are in place, a private cohort can take a baseline on your sample. Until they are, the public records are the honest exhibit: classroom, demo data, dated 2026.
Read a record, then book a call
If the protocol matches how you already work, the next step is a 25-minute scoping call. We will tell you which open class is the closest fit, or whether a private cohort is required because the sample cannot travel.