Patterns from daily work with AI coding agents
Mehdi CHIBOUNI
Week-end d'intégration ATS · October 2, 2026
Outline
01
What the agent may do, and how that's enforced
02
What it should know before it starts
03
Several agents without stepping on each other
04
From merge to prod, and anything public
05
Turning a personal setup into a shared one
Each part: the pattern, a real example, and the version worth copying.
The data
1,382
prompts on client work
162
sessions, 100 of them subagents
167
memory files across 14 repos
~90%
of tool calls are Bash
~50
characters in the median prompt
Lots of shell, lots of memory, and short follow-up prompts. The work that matters happens in the setup around the prompt.
Guardrails, knowledge, parallelism, shipping. Each one ends up as a file you can version and share.
The idea behind everything
1
Gone next session
2
Personal, may be skipped
3
Shared, still advice
4
Fails the build
5
Blocks the action
6
Can't happen
Guardrails · 1
Environments, protected namespaces, trusted domains, where secrets live.
Destructive commands are fine in general, not against production.
Better
Tune it, don't babysit it. Each manual approval is a gap in the policy: review the blocks monthly, and keep the policy in a shared repo per client.
Guardrails · 2
rm:*, curl:*, sudoFound in my own setup while preparing this talk.
Guardrails · 3
1 · Incident
Apps launched from the agent's shell inherited its environment. Transcript saving turned off, silently.
2 · Memory
A rule with its reason: launch GUI apps with env -i. Works until it isn't read.
3 · Hook
A PreToolUse guard rejects the command. Exit 2: blocked, and the model is told why.
Prompts are advice. Hooks are enforced.
Better
Hooks are code: in a repo, tested, shipped to everyone. And a hook is a guardrail, not a security boundary. Isolation comes from permissions and the sandbox.
Guardrails · 4
A hook can also run after the agent edits a file. Here: every time a Python file changes, the test suite runs. If a test fails, the agent sees the output and must fix it before moving on.
Better
Run only the tests related to the changed file, with a time limit. Run the full suite once, when the agent says it's finished.
Guardrails · 5
The agent writes the command; a human copies it and runs it with ! so the output lands in the session.
Better
A read-only identity: RBAC read-only kube context, a readonly DB role, short-lived credentials. Humans approve writes only.
--context=, never switch contextGuardrails · 6
Better
Any run without a human watching goes in a sandbox, never with permissions turned off. Review changes to hooks, CI and build files before merging.
Knowledge · 1
~470
human review comments distilled into one repo's CLAUDE.md. Committed and shared.
## Ce que la review exige
(extrait de ~470 commentaires de review humains)
- Removing a DTO field → companion PR in front-v2
- …
Invariants seulement.
Pour l'historique récent, lire git log.
Better
If a rule can be a lint rule or a test, make it one. CLAUDE.md keeps the rest, rebuilt every quarter, and the review comments on those topics should drop.
Knowledge · 2
| Memory rule | Better home | Why |
|---|---|---|
| no_local_runs_against_remote_envs | Hook | Block migrations whose DB host isn't local |
| kubectl_inline_context | Hook | Reject kubectl without --context |
| no_knip_exceptions | CI + CODEOWNERS | Config changes need a reviewer |
| crash_dont_degrade | CLAUDE.md | A design rule, not checkable |
| purge_prs_no_scope_creep | MR template | Scope stated up front |
Keep the format: the rule, a Why, a How to apply. Never secrets, never names. Re-check before relying on it.
Knowledge · 3
With 300+ memory notes, loading all of them wastes context. So the notes are grouped by topic, and a short index tells the agent when to read each group.
A custom skill searches old conversations in two steps: first the list of session titles, then the full text of the few that match.
Better
Give each note an owner and a review date. Delete or merge notes that no longer apply.
Knowledge · 4
| Repo | CLAUDE.md | copilot-instructions.md |
|---|---|---|
| admin-back | 359 lines | 9 lines |
| ecommerce-back | 20 lines | 7 lines |
| ai-ui | 53 lines | 53 lines (in sync) |
Each AI tool reads its own file, and the files drift apart. The same happened with memory: Claude's notes were copied by hand into a condensed version for Codex.
Better
Write the rules once, in AGENTS.md, and make the other files point to it (symlink or a one-line import). Same for memory: one plain Markdown source.
Parallelism · 1
_wt/ab-vat-rule-zoho-tax
admin-back
1 agent · 1 branch · 1 PR
_wt/eb-vat-rule-zoho-tax
ecommerce-back
1 agent · 1 branch · 1 PR
_wt/fv2-vat-rule-zoho-tax
front-v2
1 agent · 1 branch · 1 PR
Better
Any worktree runs locally: a bootstrap script for .env, ports and DB. Use the harness's built-in worktree isolation, and clean up on merge.
Parallelism · 2
01
Backend, frontend, pipelines and infra, in parallel
02
Sized for the sprint and the people in it
03
One fix per card, with a self-contained brief
04
Confirm → fix → test → second agent reviews
Better
Checking goes inside the brief, not after it. Each card starts by confirming the issue still exists and ends with the test that proves it's fixed.
Parallelism · 3
Cross-model review
Small scripts do it: codex-review (OpenAI), agy-review (Gemini), llmcheck (both, on any text).
Better
The reviewer runs in a read-only sandbox and returns a structured verdict: agree, disagree, why.
Handoff
Two instances in two environments; one repaired the other's broken config from a markdown note.
Better
Use the harness's session messaging, not loose files someone has to place by hand.
Parallelism · 4
Test 1
The same task, run without supervision in both. The set-up repo has a CLAUDE.md listing tools, and hooks. Results side by side.
Test 2
An output-trimming proxy vs nothing: 150 vs 152 credits, 7/7 tests passing in both. 92% of input was already cached.
Verdict on test 2: not worth installing.
Better
Run each case at least 5 times and compare averages. Run in a sandbox, not with permissions turned off. Share the test script so anyone can reuse it.
Shipping · 1
merge
staging
preprod
prod
staging→preprod→prod;
all 3 live, digest-identical
Invariant: preprod and main end byte-identical to staging.
Better
Approval lives in the platform: MR approvals, protected environments, Flux gates. The agent drives the chain; it can't approve it.
Shipping · 2
"Fix if convinced, counter-comment if not."
How incoming review comments get triaged. The agent says which, and why.
Better
Make it a rule, not a habit: a hook on gh pr comment / edit that needs a confirmed draft.
Shipping · 3
Code change
The failing test, then the passing run
Deploy
Live pods, image digest, healthy checks
Bug fix
The request that failed, now succeeding
Review
What changed, nothing else
Ask for the evidence in the brief, not a screenshot afterwards.
Steering
Rule of thumb
Short follow-ups are fine once the goal and the finish line are written down.
Setup · 1
CLAUDE_CONFIG_DIR=…
Own memory, plugins and threat model. Nothing crosses between engagements.
kubectl · gh · aws · scw
~90% of tool calls are Bash. Scope each CLI's credentials to what the agent needs.
low · high · max
High for design and debugging, lower for routine work. Effort costs time and tokens.
Setup · 2
| gh / glab | PRs, MRs, review comments |
| rg | search (354 vs 168 for grep) |
| jq / yq | read JSON and YAML output |
| kubectl / psql | inspect running systems |
The global CLAUDE.md lists ~45 tools and 35 "use X instead of Y" rules. The agent follows them about two times out of three.
Better
--json, -o json)Honest review
| What I did | What I'd do now |
|---|---|
| Approved auto-mode blocks by hand | Tune the policy monthly; a manual approval is a policy bug |
| Let secrets reach prompts and allowlists | Secret manager, prompt hook, short allowlist |
| Corrected the agent after the fact | Definition of done and evidence in the brief |
| Kept knowledge in personal memory | Promote it to hooks, CI, CLAUDE.md, shared skills |
| Typed "approved" to ship | Approvals in the platform, not the prompt |
What's next
Every guardrail, review rule and pattern so far lives in a single config. None of it has to.
Next · 1
Installed with one line in settings. Updates reach everyone.
Reviewed like code. A skill without an eval doesn't merge.
Memory and prompts stay yours. Only what generalises moves in.
Next · 2
Isolation
Own memory, plugins and threat model per client. Nothing crosses between engagements.
Guardrails
Any agent mishap becomes a short blameless write-up and an MR to hooks/.
Knowledge
Every repo mines its reviews: lint rules and tests first, CLAUDE.md for the rest.
Review
A read-only review skill on every MR. People spend their attention on design.
Harvest
Once a month, lift personal rules that apply to all into shared skills.
Learning
One session that worked, one that went wrong. Anonymised.
Next · 3
| What | How |
|---|---|
| Owners | Two rotating champions per quarter curate the repo and review MRs |
| Security baseline | Secret-scan prompt hook by default; no secrets in settings or memory |
| Measure outcomes | OpenTelemetry for usage; MR cycle time, review rounds, reverts, agent incidents for results |
| Start small | Month 1: hooks and one blueprint. Month 2: review skills. Month 3: review in CI |
First step: ats/ai/skills is open. Bring one hook, one review rule, or one pattern.
Summary
| Part | Pattern | Aim for |
|---|---|---|
| Guardrails | Threat model in auto mode; hooks from incidents | Rules low on the ladder; read-only identities; no secrets inline |
| Knowledge | CLAUDE.md from reviews; typed memory | Checks before prose; memory feeds the shared repo |
| Parallelism | Worktree per feature per repo; audit → agents | Worktree-ready envs; checking inside each brief |
| Shipping | Agent drives the release; drafts before posting | Approvals in the platform; evidence before done |
| Next | ats/ai/skills | One contribution each |
« Move every rule as far down the ladder as it will go. »
Happy to share: the hook, the auto-mode block, and a CLAUDE.md skeleton built from reviews.