ATS Digital Dev

Working with agents

Patterns from daily work with AI coding agents

Mehdi CHIBOUNI

Week-end d'intégration ATS · October 2, 2026

Outline

Five parts

01

Guardrails

What the agent may do, and how that's enforced

02

Knowledge

What it should know before it starts

03

Parallelism

Several agents without stepping on each other

04

Shipping

From merge to prod, and anything public

05

What's next

Turning a personal setup into a shared one

Each part: the pattern, a real example, and the version worth copying.

The data

The setup, in numbers

1,382

prompts on client work

162

sessions, 100 of them subagents

167

memory files across 14 repos

~90%

of tool calls are Bash

~50

characters in the median prompt

Lots of shell, lots of memory, and short follow-up prompts. The work that matters happens in the setup around the prompt.

The engineering goes into the harness, not the prompt.

Guardrails, knowledge, parallelism, shipping. Each one ends up as a file you can version and share.

The idea behind everything

Move each rule down the ladder

1

Chat

Gone next session

2

Memory

Personal, may be skipped

3

CLAUDE.md

Shared, still advice

4

Lint / test

Fails the build

5

Hook

Blocks the action

6

Permission / sandbox

Can't happen

Guardrails · 1

Auto mode reads a written threat model

~25 lines of context

Environments, protected namespaces, trusted domains, where secrets live.

Prod-only soft-denies

Destructive commands are fine in general, not against production.

Better

Tune it, don't babysit it. Each manual approval is a gap in the policy: review the blocks monthly, and keep the policy in a shared repo per client.

"autoMode": { "environment": [ "Protected envs: production2/*", "Trusted domains: *.<internal>", "Secrets: ExternalSecret manifests", … ~25 lines ], "soft_deny": [ "$defaults", "kubectl delete in production2 namespaces", "aws s3 rm targeting the prod DWH bucket" ] }

Guardrails · 2

Allowlists and transcripts collect secrets

What piles up

  • "Always allow" saves the full command, inline passwords included
  • Broad rules creep in: rm:*, curl:*, sudo
  • A key pasted in a prompt stays in a transcript on disk

Instead

  • Credentials come from env or a secret manager, never inline
  • Short allowlist; auto mode and hooks cover the rest
  • A prompt hook that blocks anything shaped like a key
  • Review the allowlist before sharing a config

Found in my own setup while preparing this talk.

Guardrails · 3

From incident to memory to hook

1 · Incident

Apps launched from the agent's shell inherited its environment. Transcript saving turned off, silently.

→

2 · Memory

A rule with its reason: launch GUI apps with env -i. Works until it isn't read.

→

3 · Hook

A PreToolUse guard rejects the command. Exit 2: blocked, and the model is told why.

Prompts are advice. Hooks are enforced.

Better

Hooks are code: in a repo, tested, shipped to everyone. And a hook is a guardrail, not a security boundary. Isolation comes from permissions and the sandbox.

Guardrails · 4

Run the tests after every edit

A hook can also run after the agent edits a file. Here: every time a Python file changes, the test suite runs. If a test fails, the agent sees the output and must fix it before moving on.

Better

Run only the tests related to the changed file, with a time limit. Run the full suite once, when the agent says it's finished.

.claude/settings.json "PostToolUse": [{ "matcher": "Edit|Write", "hooks": [ { "command": "hooks/shellcheck.sh" }, { "command": "hooks/pytest.sh" } ] }] # pytest.sh: tests fail → exit 2 # → the agent reads the failure

Guardrails · 5

Give the agent a read-only way into prod

Today

The agent writes the command; a human copies it and runs it with ! so the output lands in the session.

Better

A read-only identity: RBAC read-only kube context, a readonly DB role, short-lived credentials. Humans approve writes only.

Rules that hold either way

  • Always pass --context=, never switch context
  • No port-forwards started from the session
  • No polling loops against the prod DB
  • Never run locally against a remote env

Guardrails · 6

Unattended agents run in a sandbox

What the sandbox gives you

  • The agent runs in a small VM, not on your machine
  • API keys are added by a proxy on the host: the real values never enter the VM
  • The repo can be mounted read-only; the agent works on a copy

What it doesn't protect

  • Inside the VM, the agent has full sudo
  • Local MCP servers still run on your machine
  • Files that run later on your machine: git hooks, CI config, Makefile, package.json scripts

Better

Any run without a human watching goes in a sandbox, never with permissions turned off. Review changes to hooks, CI and build files before merging.

Knowledge · 1

CLAUDE.md, mined from our PR reviews

~470

human review comments distilled into one repo's CLAUDE.md. Committed and shared.

## Ce que la review exige

(extrait de ~470 commentaires de review humains)

- Removing a DTO field → companion PR in front-v2
- …


Invariants seulement.
Pour l'historique récent, lire git log.

Better

If a rule can be a lint rule or a test, make it one. CLAUDE.md keeps the rest, rebuilt every quarter, and the review comments on those topics should drop.

Knowledge · 2

Memory is a staging area, not the destination

Memory ruleBetter homeWhy
no_local_runs_against_remote_envsHookBlock migrations whose DB host isn't local
kubectl_inline_contextHookReject kubectl without --context
no_knip_exceptionsCI + CODEOWNERSConfig changes need a reviewer
crash_dont_degradeCLAUDE.mdA design rule, not checkable
purge_prs_no_scope_creepMR templateScope stated up front

Keep the format: the rule, a Why, a How to apply. Never secrets, never names. Re-check before relying on it.

Knowledge · 3

Load memory only when it's relevant

With 300+ memory notes, loading all of them wastes context. So the notes are grouped by topic, and a short index tells the agent when to read each group.

MEMORY.md - index_work.md LOAD THIS whenever touching GitLab, Kubernetes, … - index_system.md - …

Finding past sessions

A custom skill searches old conversations in two steps: first the list of session titles, then the full text of the few that match.

Better

Give each note an owner and a review date. Delete or merge notes that no longer apply.

Knowledge · 4

One instruction file for every AI tool

RepoCLAUDE.mdcopilot-instructions.md
admin-back359 lines9 lines
ecommerce-back20 lines7 lines
ai-ui53 lines53 lines (in sync)

Each AI tool reads its own file, and the files drift apart. The same happened with memory: Claude's notes were copied by hand into a condensed version for Codex.

Better

Write the rules once, in AGENTS.md, and make the other files point to it (symlink or a one-line import). Same for memory: one plain Markdown source.

Parallelism · 1

One worktree per feature, per repo

feature: vat-rule-zoho-tax

_wt/ab-vat-rule-zoho-tax

admin-back

1 agent · 1 branch · 1 PR

_wt/eb-vat-rule-zoho-tax

ecommerce-back

1 agent · 1 branch · 1 PR

_wt/fv2-vat-rule-zoho-tax

front-v2

1 agent · 1 branch · 1 PR

Better

Any worktree runs locally: a bootstrap script for .env, ports and DB. Use the harness's built-in worktree isolation, and clean up on merge.

Parallelism · 2

Audit, then cards, then agents

01

3 read-only agents

Backend, frontend, pipelines and infra, in parallel

→

02

Ranked findings

Sized for the sprint and the people in it

→

03

Cards

One fix per card, with a self-contained brief

→

04

1 agent per card

Confirm → fix → test → second agent reviews

Better

Checking goes inside the brief, not after it. Each card starts by confirming the issue still exists and ends with the test that proves it's fixed.

Parallelism · 3

Agents checking agents

Cross-model review

A different model gets a second look

Small scripts do it: codex-review (OpenAI), agy-review (Gemini), llmcheck (both, on any text).

Better

The reviewer runs in a read-only sandbox and returns a structured verdict: agree, disagree, why.

Handoff

Agents leave notes for each other

Two instances in two environments; one repaired the other's broken config from a markdown note.

Better

Use the harness's session messaging, not loose files someone has to place by hand.

Parallelism · 4

Measure an AI setup before adopting it

Test 1

Plain repo vs set-up repo

The same task, run without supervision in both. The set-up repo has a CLAUDE.md listing tools, and hooks. Results side by side.

Test 2

Token-saving tools

An output-trimming proxy vs nothing: 150 vs 152 credits, 7/7 tests passing in both. 92% of input was already cached.

Verdict on test 2: not worth installing.

Better

Run each case at least 5 times and compare averages. Run in a sandbox, not with permissions turned off. Share the test script so anyone can reuse it.

Shipping · 1

The agent drives the release chain

merge

→

staging

→

preprod

→

prod

What comes back

staging→preprod→prod;
all 3 live, digest-identical

Invariant: preprod and main end byte-identical to staging.

Better

Approval lives in the platform: MR approvals, protected environments, Flux gates. The agent drives the chain; it can't approve it.

Shipping · 2

Draft first, post second

plans/…md What's left is only on GitHub (outward-facing, so planned here first). 1. gh pr edit … --body-file … 2. Update the companion PR description 3. Reply on the reviewer's thread Verify: gh pr view --json body

"Fix if convinced, counter-comment if not."

How incoming review comments get triaged. The agent says which, and why.

Better

Make it a rule, not a habit: a hook on gh pr comment / edit that needs a confirmed draft.

Shipping · 3

Done means evidence

Code change

Test output

The failing test, then the passing run

Deploy

Rollout status

Live pods, image digest, healthy checks

Bug fix

The real response

The request that failed, now succeeding

Review

The diff

What changed, nothing else

Ask for the evidence in the brief, not a screenshot afterwards.

Steering

Put the effort in the first prompt

A first prompt that holds up

Goal: what should change, and why Context: ticket, files, related PRs Constraints: what must not change Done when: the test passes, the rollout is healthy, the response is 200 Stop and ask: before migrations, prod, anything public

What to avoid

  • "carry on" with no definition of done: the agent decides when it's finished
  • "commit push pr" without reading the diff: you ship what you haven't reviewed
  • Asking for "tldr" every time: say it once in CLAUDE.md

Rule of thumb

Short follow-ups are fine once the goal and the finish line are written down.

Setup · 1

Configure the harness per context

One config per client

CLAUDE_CONFIG_DIR=…

Own memory, plugins and threat model. Nothing crosses between engagements.

The terminal is the interface

kubectl · gh · aws · scw

~90% of tool calls are Bash. Scope each CLI's credentials to what the agent needs.

Effort per task

low · high · max

High for design and debugging, lower for routine work. Effort costs time and tokens.

Setup · 2

Tell the agent which command-line tools to use

What the agent actually runs

gh / glabPRs, MRs, review comments
rgsearch (354 vs 168 for grep)
jq / yqread JSON and YAML output
kubectl / psqlinspect running systems

The global CLAUDE.md lists ~45 tools and 35 "use X instead of Y" rules. The agent follows them about two times out of three.

Better

  • A short list per project, not 45 tools
  • Ask for JSON output (--json, -o json)
  • Tokens limited to what the agent needs
  • Same tools for everyone: a Brewfile in the project template
  • A hook that swaps grep for rg, instead of a rule it may ignore

Honest review

What I'd do differently

What I didWhat I'd do now
Approved auto-mode blocks by handTune the policy monthly; a manual approval is a policy bug
Let secrets reach prompts and allowlistsSecret manager, prompt hook, short allowlist
Corrected the agent after the factDefinition of done and evidence in the brief
Kept knowledge in personal memoryPromote it to hooks, CI, CLAUDE.md, shared skills
Typed "approved" to shipApprovals in the platform, not the prompt

What's next

What if the harness were shared?

Every guardrail, review rule and pattern so far lives in a single config. None of it has to.

Next · 1

One shared repo: ats/ai/skills

gitlab.ats-digital.com/ats/ai/skills ├── blueprints/ new project: CLAUDE.md, settings, CI ├── reviews/ review skills per stack, from our MRs ├── patterns/ worktrees, audit → cards → agents ├── hooks/ prod guard, secret guard, env guard ├── automode/ threat-model templates per client ├── agents/ briefs with evidence built in └── evals/ every skill ships with a test

A plugin marketplace

Installed with one line in settings. Updates reach everyone.

Contribute by MR

Reviewed like code. A skill without an eval doesn't merge.

Personal stays personal

Memory and prompts stay yours. Only what generalises moves in.

Next · 2

Where it could go

Isolation

One config per client

Own memory, plugins and threat model per client. Nothing crosses between engagements.

Guardrails

Incident → shared hook

Any agent mishap becomes a short blameless write-up and an MR to hooks/.

Knowledge

Checks before prose

Every repo mines its reviews: lint rules and tests first, CLAUDE.md for the rest.

Review

Agent review in CI

A read-only review skill on every MR. People spend their attention on design.

Harvest

Promote memories

Once a month, lift personal rules that apply to all into shared skills.

Learning

Session show-and-tell

One session that worked, one that went wrong. Anonymised.

Next · 3

Making it stick

WhatHow
OwnersTwo rotating champions per quarter curate the repo and review MRs
Security baselineSecret-scan prompt hook by default; no secrets in settings or memory
Measure outcomesOpenTelemetry for usage; MR cycle time, review rounds, reverts, agent incidents for results
Start smallMonth 1: hooks and one blueprint. Month 2: review skills. Month 3: review in CI

First step: ats/ai/skills is open. Bring one hook, one review rule, or one pattern.

Summary

The patterns, in their better form

PartPatternAim for
GuardrailsThreat model in auto mode; hooks from incidentsRules low on the ladder; read-only identities; no secrets inline
KnowledgeCLAUDE.md from reviews; typed memoryChecks before prose; memory feeds the shared repo
ParallelismWorktree per feature per repo; audit → agentsWorktree-ready envs; checking inside each brief
ShippingAgent drives the release; drafts before postingApprovals in the platform; evidence before done
Nextats/ai/skillsOne contribution each

« Move every rule as far down the ladder as it will go. »

Merci · Questions?

Happy to share: the hook, the auto-mode block, and a CLAUDE.md skeleton built from reviews.