Server Plugin: Please Make Quarantine Read-Only and Improve 4429 documentation
I want to report a fairly serious issue I ran into with the <skill_content name="server-plugin"> workflow and suggest some changes that would make server failures much easier to diagnose and much safer to recover from.
My server became stuck in quarantine, and it was extremely difficult to determine what was actually causing it. Even AI helper could not reliably identify the underlying problem.
The particularly problematic part was that tests and server-code/version updates could repeatedly trigger quarantine, and each time I was effectively locked out of the server for around 60 minutes. This made debugging extremely difficult because every attempted fix could result in another quarantine period.
I think the server plugin and especially the AI-helper guidance should handle this situation differently.
Suggestions
1. Make quarantine read-only instead of completely inaccessible
When the server enters quarantine, the server should ideally become read-only rather than being completely locked out.
Writes/mutations could remain disabled until the quarantine is resolved.
This would make quarantine a safe mode rather than a complete black box.
2. Treat WebSocket close code 4429 differently from ordinary network failures
The AI-helper/server-plugin documentation should explicitly explain what 4429 means and how an agent should react to it.
It should not be treated like:
"The connection failed, reconnect normally."
A 4429 quarantine is fundamentally different from an ordinary temporary network failure.
If an AI agent receives 4429, it should stop repeatedly modifying/testing/redeploying the server blindly.
Instead, it should enter a diagnostic workflow:
-
Detect
4429. -
Stop aggressive reconnect/retry behavior.
-
Determine whether the server is in quarantine.
-
Read available diagnostic information.
-
Inspect storage/snapshot state if read access is available.
-
Only after the cause is understood should the agent propose a write, migration, repair, or server-code change.
-
Identify whether the problem is:
- corrupted data,
- malformed record,
- sequence mismatch,
- invalid snapshot,
- migration problem,
- server-code problem,
- storage-size problem,
- or an actual temporary service/network condition.
This would prevent an AI helper from getting into a loop like:
change server → deploy → 4429 → wait → test → change server → deploy → 4429 → wait again
That loop is extremely expensive when each quarantine can lock the developer out for approximately an hour.
3. The AI-helper skill should have an explicit 4429 recovery procedure
I think <skill_content name="server-plugin"> should contain a dedicated section such as:
### 4429 Quarantine Recovery
Do NOT treat 4429 as an ordinary network failure.
When 4429 occurs:
* stop repeated reconnect attempts;
* do not repeatedly deploy modified server code just to test whether the problem disappeared;
* do not perform destructive migrations;
* use read-only diagnostics first;
* inspect storage metadata and snapshot validation results;
* identify the exact failure before making changes;
* preserve/export the affected storage state where possible;
* only perform repair/migration after the corruption/failure has been identified;
* respect the server-provided quarantine/retry interval.
If diagnostic read access is available while quarantined, use it before attempting another deployment or mutation.
That would give an AI coding agent a much safer operating procedure.
4. Distinguish three different situations
The skill documentation should clearly distinguish:
Normal network failure
connection lost
→ reconnect with normal backoff
Temporary server/service problem
temporary server unavailable / overload
→ respect retry information
→ reconnect later
4429 quarantine
server entered quarantine
→ STOP normal retry/debug loop
→ enter diagnostic/read-only mode
→ inspect server state
→ determine cause
→ repair only after diagnosis
Those should not all be handled by the same generic reconnect logic.
5. Preserve diagnostic access even when writes are disabled
I understand why the server might need to prevent writes during quarantine. What caused the biggest problem was losing the ability to inspect what was actually happening.
A good compromise would be:
Quarantine = no mutations, but diagnostic reads remain available.
That would protect the server from further corruption while still allowing the developer and AI-helper to investigate it.
Main request
I would strongly recommend improving <skill_content name="server-plugin"> so that quarantine is treated as a diagnostic/recovery state rather than a complete lockout, and adding a dedicated, explicit 4429 procedure for AI-helper agents.
The biggest practical problem I experienced was not simply that corruption/quarantine happened. It was that once it happened, diagnosing the cause became extremely difficult, and every attempted test or version change could trigger another long quarantine cycle.
A read-only quarantine mode plus detailed 4429 documentation for AI agents would make these failures substantially easier and safer to troubleshoot without repeatedly locking developers out for long periods.
1 reply
And please tell the AI helper don't wipe out server data for upgrading a version, thank you dev