← back to journal
EngineeringJun 12, 20264 min

The AI length-checked a dropdown — and people call this AGI

arc: a 'top-tier' ai length-checks a 3-option dropdown, then burns 100k tokens answering the wrong question → it doesn't infer intent, it just pattern-matches → it's a tool, not a mind. validate everything.

small bug in faculty enrollment. gender is mandatory, but the form let you skip it and move on anyway 🤷‍♂️

i asked claude to fix it. it patched the frontend, then said the backend needed a guard too — fair, the data layer had no check at all. then i read what it actually wrote.

the AI's rule:
  body('gender').isString().notEmpty().isLength({ max: 20 })

the field on screen:
  a dropdown.  Male / Female / Other.  three options. that's it.

a length check. on a dropdown 🥲

isLength({ max: 20 }) happily accepts "banana". it accepts any twenty characters of garbage. it was never guarding gender — it was guarding the number of letters in a field that can only ever be one of three exact words. the correct rule isn't subtle:

  body('gender').isIn(['Male', 'Female', 'Other'])

required and valid-value, in one line. the code it wrote was valid — it runs, it catches empty. it just wasn't correct. only knowing the domain told me which.

Claude's diff in teacherValidator.ts replacing the gender length check with an isIn allowlist of Male, Female, Other

the length check (red) becomes the allowlist (green) — one line, and it's actually correct.

so before correcting it, i asked it straight: are you doing this on purpose, or do you genuinely not notice you're counting characters on a three-option dropdown?

it said sorry. no edge case it was guarding, no reasoning to defend — because there was never any thought there. it counted letters on a closed dropdown, and the second i pointed, it folded.

i asked for the root cause. its own word: pattern-mimicry. it had grabbed the nearest validator that touched gender (studentValidator.ts), seen isLength({ max: 20 }), and copied the shape of it — without ever asking what a thinking thing asks itself: wait, this is a dropdown. why am i measuring its length?

Claude's own root-cause write-up calling the mistake pattern-mimicry — it copied the shape of the neighbouring student validator instead of reasoning from what the gender field actually is

its own post-mortem, word for word — read it as a confession.

know what the field actually is → write the rule that fits it → let the DB column be a backstop, never the source of truth.

and because it copies instead of thinks, it spreads the mistake. if it cloned the student validator, the student validator had the same bug — same length check, sitting in production. so i sent it back to fix that one too.

Claude confirming gender is a closed dropdown so the correct rule is an isIn allowlist, and noting studentValidator.ts still carried the same looser length check

it even admitted the student validator had the same weakness — the pattern it cloned was wrong at the source.

then it got worse. i asked it to sweep the whole codebase for the same class of problem — anywhere the frontend and backend disagree on what a field is allowed to be. it spent 100k tokens crawling the repo and came back with: "no other issues with gender dropdowns." 🤦‍♂️

i never asked about gender dropdowns. that was one example. i asked about validation mismatches, everywhere. it took my literal words and optimised for the narrowest reading of them — burning a fortune in tokens to answer a question i didn't ask.

people will say i prompted it badly. fine — but look at what that defence actually concedes:

Ex: if it needs me to spell out every word, then it isn't doing the thinking — i am. understanding that "similar issues" means the pattern, not the one noun, is the easy part of the job. that's the part it failed.

a junior gets this on the first try — say "check for similar issues" and they know similar means the pattern. this didn't, and not for lack of effort: it spent 100k tokens. it just doesn't grasp intent. it matches strings.

and here's what gets me: this is claude 4.8, a "top-tier" model, and people queue up to tell me it's nearly AGI and my job is in danger. it can't tell that a three-option dropdown doesn't need its length measured, and it can't tell that "similar issues" means the pattern. every model i've used is the same underneath — fast, confident, pattern-matching from whatever's nearest. they don't think, and they don't understand your product.

so, the unhyped version: it does not replace a junior. the companies selling the AGI story are mostly selling themselves. the tool is genuinely useful — it is not a mind, and pretending it is one is how dumb bugs reach production.

use it like the tool it is. read every line like a junior's PR — the only reason any of this got caught is a habit my team lead Jagan Kumar Mudila sir drilled into me: never trust AI code because it looks right or because it compiles. validate everything. the understanding is still your job, and it isn't going anywhere 🤷‍♂️

#buildinpublic #softwareengineering #ai #maahitatechnologies

Send this as proof →Share on LinkedIn