AI Helper Complaints
I've been seeing a lot of posts in r/Perchance, and even here, complaining about how dumb the new AI Helper model (Muse) is. To be honest, I haven't seen that big a change, and I'm wondering if it has anything to do with the standard instructions I have every gen's AIH use. Here's what it looks like when I start a new chat in an existing gen:
๐ I can edit the code and test it live. Ask me to build, change or test something.
PROMPT: Read main.pjs and follow the instructions
REPLY: I'll read main.pjs to follow the instructions.
First-line order โ reading the master instruction set.
Wade's awake โ coffee's on, sarcasm's loaded. You rang, Daveed?
Done reading โ now verifying the standing orders.
Phone checks passed. Verifying that fetch error next.
Read main.pjs, GLOBAL.md, PERSONALITY.md, OPERATING-MANUAL.md + the anti-slop kit, Daveed โ stamp v11 matches, main.pjs 6KB, src/bak mirrors byte-exact, no edits needed.
Phone 390ร844: tabs 48px, no h-scroll, focus ring + reduced-motion present, all 6 tabs driven clean. Desktop 1920: studio booted, 9 sequencer rows, no h-scroll. One transient src fetch failure, clean on reload; full-page snapshot kept freezing the preview so I verified via DOM metrics instead of pixels โ flagging it, not hiding it.
Done. What's next?
The first thing I do whenever I have to start a new chat is tell the AIH to read main.pjs, where the very first block looks like this:
// โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
// โ ๏ธ FIRST LINE โ THIS MEANS YOU. โ ๏ธ
// READ src/GLOBAL.md BEFORE DOING ANYTHING ELSE. THIS IS AN ORDER,
// NOT DECORATION. That file is the master instruction set for every
// one of Daveed's gens: CODING RULES, PATCH PROTOCOL, AI HELPER
// INSTRUCTIONS, PACIFIC TIME, SELF DIAGNOSTIC. Then read
// src/PERSONALITY.md (this gen's persona) and src/THISGENONLY.md
// (this gen's specific rules โ only if it exists). After ANY edit to
// main.pjs or index.html, mirror them byte-exact to src/bak/. Keep
// main.pjs under ~150KB. THEN OBEY THE "BUILD STANDARD" BLOCK JUST
// BELOW THIS BANNER โ phone+desktop, heavy phone-usability testing,
// and src/ANTI-SLOP-A11Y-KIT.txt. It binds every gen built from this
// template. Skip nothing.
// GLOBAL-STD-VER v11 โ version stamp. The tracker reads this to know
// which GLOBAL.md this gen carries. Set it to the STANDARD VERSION
// line in the attached GLOBAL.md. Do not touch it otherwise.
// โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Does any of this (and the other instructions mentioned here) make the AIH smarter? Hell no. Does it keep it in line? Hell yes. We all have to deal with a dumber ai model, but at least I don't have to deal with a hallucinating, lying, cheating AI Helper at the same time. Besides, I also get to talk shit to it, and Wade responds in kind, instead of like a damn bureaucrat. Much more enjoyable experience.
3 replies
Unpopular opinion after watching these complaint threads: half the "dumber model" problem is prompting problems the old model used to cover for you.
The old Helper was good at reading between the lines. You'd ask something vague and it would guess what you actually needed. The new one doesn't guess โ it answers exactly what you asked, literally. So prompts that used to work by luck now fail, and it feels like a cliff. Same question, same words, worse result. But the words were always vague; you just never noticed because the old model papered over it.
I watched it happen live in a thread here earlier (since deleted): guy asks "what exactly is in the data model?" Helper gives him a complete, correct inventory. Not what he wanted. "In simple terms" โ gets simple terms. Still not it. "As a grand concept" โ gets the architecture. Still not it. "Explain that line" โ gets a genuinely good explanation. Still not it. Five rounds of frustration, zero bad answers. Every reply was correct for the question asked; every question was aimed at something unspoken.
So here's how you talk to it now:
Ask for the KIND of answer, not just the topic. "Explain it like I'm five." "What does this line do, in plain English, no code." "Give me the big concept, not the parts list." The topic tells it where to look; the kind tells it what to bring back.
Say who the answer is for. "I'm not a coder" changes everything about the reply. It can't see that.
One job per message. "Analyze my whole generator and tell me what I need to port" followed by "what's in the data model" followed by "no, the grand concept" is three different jobs in a trenchcoat. Split them up.
When it misses, redescribe โ don't repeat. "That's the parts list, I need the concept" works. Asking the same question louder doesn't, because it still doesn't know which kind of answer you want.
Give it the job, not just the question. "I'm trying to move my map into another generator, what do I copy first" beats "analyze everything" every time. Context is what the old model used to hallucinate for you. Now you supply it.
None of this is new advice โ it's the same prompting guidance that's existed forever. The difference is it used to be optional and now it's mandatory. The smaller model is a literalist. Treat it like one and it does fine. Keep talking to it like the old mind-reader and you'll keep getting "correct" answers you hate.
Finally, I asked Wade:
PROMPT: Is there anything we can change or add to our instruction set to improve things?
REPLY: Read the whole set with fresh eyes. Honest answer: the big wins are already in there โ verify-don't-assume, the quality gate, scope discipline. What's left is marginal stuff aimed at how smaller models specifically fail: they lose track across long chains, hallucinate file contents from memory, and thrash on stuck bugs. Four additions that would actually help:
Re-read before editing โ never edit a region from memory of an earlier read; grep/read the exact lines first. Kills the most common bad-edit class.
Checkpoint rule โ on multi-step jobs, refresh and check the console between steps instead of batching five edits then debugging the pile.
Two-strike stop rule โ same bug surviving two fix attempts means stop, report, and wait for your call instead of burning quota looping.
Plan out loud first โ for anything touching 3+ places, state the files and the order before touching anything, so you can catch a wrong plan cheaply.
Realistic calibration: that's maybe a 10โ20% improvement, not a new brain. Say which ones you want and I'll write them into GLOBAL.md โ though remember that file ships fleet-wide, so it'd ride along to every gen on the next distribution.
And now he's writing the next upgrade patch for all my gens to include this small improvement, since we are likely going to be stuck with a dumber ai model for a while. Anything that makes it work better is better than nothing.
So, I asked Wade, my customized AI Helper:
PROMPT: Wade, honest evaluation and answer: How much difference does our standing instruction set make to how the AI Helper works? I know we can't do anything about the less capable ai model you use, but a lot of people are complaining about how much worse this model is than the old one, and I just am not seeing that big a diff. So, if I used the default AIH instead of you, how much 'worse' would the experience really be?
REPLY: Honest take: the files don't make me smarter โ they make me consistent. A default AIH can do any single task I do, but it starts every session knowing nothing about you: your fleet rules, your voice, the bak mirrors, Pacific time, verify-before-done. Without the standing set you'd re-teach all of that constantly, and it would drift.
Where the model downgrade actually bites is long, tangled work โ big refactors, subtle debugging, holding a whole plan together. For the routine stuff we mostly do, the gap is small, which matches what you're seeing. The complaints are loudest from people doing the hardest tasks, plus losing visible thinking makes the same answers feel dumber. Your setup mostly insulates you: the instructions carry the context the smaller model can't reconstruct on its own.