orchestration, prompts, evals, protocols, and the LLM wallet
agent-stack
Production patterns for AI agent systems, in four skills. The orchestrator: tool-calling loops that survive context pressure, pipelines with human checkpoints and resume, provider routing with fallback and health checks, four memory layers with confidence decay, and the shape of the work as a graph — an edge that carries no data is no edge, and a parallel layer gets a checker before the node that consumes it. The harness: what the agent is TOLD — system prompts, tool descriptions, technique choice, workflow versus agent, static versus dynamic, plus a seven-track audit of somebody else's agent and a scanner for the defects that are mechanically visible. The evals: suites that measure whether it actually works. The interop layer: MCP 2026-07-28 and what it deprecated, A2A 1.0, the registry, the gateway. Plus the wallet side of reselling LLM access.
Install this pack on its own, or use it with the ssheleg harness.
What it ships
Each name below is an entry point an agent can be routed to.
Shape: static. Reference material, not a run. What varies is which reference a task loads, which the SKILL.md table decides in advance.
agent-orchestrator
Use when building an agent system — an orchestrator, an LLM-powered tool, a chatbot with tool use, an AI pipeline, or metering and billing the LLM access it burns. Covers tool-calling loops, pipelines with human checkpoints, provider routing with fallback/retry, memory architecture, retrieval and decay, context budgets, sub-agent coordination, error hierarchies; the work as a graph — parallel layers, fake edges, a checker before convergence; for resale: tiered wallets, one markup boundary, the saga across database and provider API, spend-delta polling, budget and loop guards, per-tenant keys. Not for a single LLM call in a script, or prompt wording.
agent-evals
Use when measuring whether an agent actually works — building an eval suite, judging a trajectory rather than a final answer, turning production traces into regression fixtures, calibrating an LLM judge against human labels, or gating a release on offline evals. Covers the three observability primitives (run, trace, thread) crossed with three eval granularities (single-step, full-turn, multi-turn), the offline/online/ad-hoc timing axis, pass-fail rubrics over scalar scores, cheap code checks before model judges, the checker node as an evaluator inside the graph, simulated users with adversarial personas, annotation queues, and what to instrument for any of it. Not for unit tests of ordinary code, or benchmarking a model.
agent-interop
Use when an agent must talk to something outside its own process — building or consuming an MCP server, exposing or calling an agent over A2A, publishing to the MCP Registry, or putting a gateway in front of agent traffic. Carries the MCP 2026-07-28 wire surface and what it deprecated (server/discover, stateless per-request _meta, elicitation in form and URL mode, subscriptions/listen; sampling, roots, logging and dynamic client registration going), A2A 1.0 agent cards, task states, three bindings, registry namespaces, server.json, tool federation, and what a gateway must do that an API gateway does not. Not for designing one server's tool set, nor for a skill's own construction — that is make-skill.
agent-harness
Use when the question is what the agent is TOLD rather than how its loop is wired — writing or fixing a system prompt, shaping tools so the model picks the right one, deciding whether a job wants a workflow or an agent, or choosing between ReAct, reflection and voting. Also auditing an agent system somebody else built: tracks, evidence tiers and a prioritized plan instead of a score, plus a scanner. Carries Pi as a worked kernel implementation — SDK, RPC and extension seams, for embedding or extending a harness. Not for the loop's plumbing, its evals, or its protocols — those are siblings.
Install just this one
Every pack installs standalone. The whole family is one command.
npx skills add ssheleg/agent-stack
claude plugin marketplace add ssheleg/agent-stack && claude plugin install agent-stack@agent-stack
When the agent reaches for it
These are the rules the family writes into your agent's own instruction file — verbatim. Each one states the rule, the boundary in both directions, and the phrase that declines it.
/agent-stack — how an agent system is built, judged and metered — and the wire it speaks outward
- when: the thing being built is an agent, or a server one connects to
- decline it: no agent layer
If agent-stack is installed, building an agent SYSTEM goes through it — loops/work-graphs, system prompts/tools, workflow vs agent, evals, MCP/A2A, registry/gateway, and the wallet behind resold LLM access.
The boundary. NOT through it: one LLM call in a script, one prompt's wording, person-facing UI (super-ux, sheleg-design), charging money (sheleg-dev), agents editing THIS repo (agent-sync). sheleg-dev charges; agent-stack meters LLM use.
Refusal phrase: "no agent layer" or «без агентного слоя».
Among the routers: agent-sync coordinates file ownership; agent-stack owns the agent system being built.
The rest of the family
super-ux v0.59.1
Scenario-driven UI development: a versioned design chain in docs/ux/ — the product vision, personas and jobs, user flows and the paid-acquisition funnel among them, a screens-and-states map with Figma frames…
task-pipeline v1.90.2
Full-cycle delivery orchestrator: an intake grill turns the request into a complete brief, then ten gated stages carry it from docs to acceptance, refusing to advance until each gate passes.
agent-sync v1.21.5
Coordination plane for concurrent agents: leases with a TTL so two agents cannot claim the same work, race-free id reservation, a run journal and a generated board, over a pluggable knowledge cloud.
make-skill v0.29.2
A skill that builds skills: create, retrofit, audit and publish agent skills and Claude Code plugins
sheleg-design v1.64.0
The taste layer: cinematic scroll-driven landing pages (one scroll clock, motion that degrades to calm, WebGL particle formations), product-UI style packs with a ready token layer each, and the Figma border
seo-aeo-audit v0.26.3
Evidence-first website audit for search and answer engines: ten tracks from crawl access to AI citation mechanics, every finding backed by an observation and every recommendation tiered, ending in…
sheleg-dev v0.13.2
The integration layer a product reaches once it has users: Stripe subscription billing reconciled into your own database
telegram-dev v0.2.2
Telegram, split by the API each surface actually speaks, in three skills. telegram-bots covers the official HTTP Bot API
xr-dev v0.3.2
Quest product delivery from platform choice to launch and operation, with stage owners, evidence gates and next actions.
web3d-dev v0.1.2
Realtime 3D on the web with three.js and React Three Fiber, in three skills split by the question. web3d-runtime owns how the scene runs