ssheleg skills → agent-stack

orchestration, prompts, evals, protocols, and the LLM wallet

agent-stack

Production patterns for AI agent systems, in four skills. The orchestrator: tool-calling loops that survive context pressure, pipelines with human checkpoints and resume, provider routing with fallback and health checks, four memory layers with confidence decay, and the shape of the work as a graph — an edge that carries no data is no edge, and a parallel layer gets a checker before the node that consumes it. The harness: what the agent is TOLD — system prompts, tool descriptions, technique choice, workflow versus agent, static versus dynamic, plus a seven-track audit of somebody else's agent and a scanner for the defects that are mechanically visible. The evals: suites that measure whether it actually works. The interop layer: MCP 2026-07-28 and what it deprecated, A2A 1.0, the registry, the gateway. Plus the wallet side of reselling LLM access.

Install this pack on its own, or use it with the ssheleg harness.

v0.25.5pinned in this release of the family
4skills in the pack
1routing rule it owns

What it ships

Each name below is an entry point an agent can be routed to.

Shape: static. Reference material, not a run. What varies is which reference a task loads, which the SKILL.md table decides in advance.

agent-orchestrator

Use when building an agent system — an orchestrator, an LLM-powered tool, a chatbot with tool use, an AI pipeline, or metering and billing the LLM access it burns. Covers tool-calling loops, pipelines with human checkpoints, provider routing with fallback/retry, memory architecture, retrieval and decay, context budgets, sub-agent coordination, error hierarchies; the work as a graph — parallel layers, fake edges, a checker before convergence; for resale: tiered wallets, one markup boundary, the saga across database and provider API, spend-delta polling, budget and loop guards, per-tenant keys. Not for a single LLM call in a script, or prompt wording.

agent-evals

Use when measuring whether an agent actually works — building an eval suite, judging a trajectory rather than a final answer, turning production traces into regression fixtures, calibrating an LLM judge against human labels, or gating a release on offline evals. Covers the three observability primitives (run, trace, thread) crossed with three eval granularities (single-step, full-turn, multi-turn), the offline/online/ad-hoc timing axis, pass-fail rubrics over scalar scores, cheap code checks before model judges, the checker node as an evaluator inside the graph, simulated users with adversarial personas, annotation queues, and what to instrument for any of it. Not for unit tests of ordinary code, or benchmarking a model.

agent-interop

Use when an agent must talk to something outside its own process — building or consuming an MCP server, exposing or calling an agent over A2A, publishing to the MCP Registry, or putting a gateway in front of agent traffic. Carries the MCP 2026-07-28 wire surface and what it deprecated (server/discover, stateless per-request _meta, elicitation in form and URL mode, subscriptions/listen; sampling, roots, logging and dynamic client registration going), A2A 1.0 agent cards, task states, three bindings, registry namespaces, server.json, tool federation, and what a gateway must do that an API gateway does not. Not for designing one server's tool set, nor for a skill's own construction — that is make-skill.

agent-harness

Use when the question is what the agent is TOLD rather than how its loop is wired — writing or fixing a system prompt, shaping tools so the model picks the right one, deciding whether a job wants a workflow or an agent, or choosing between ReAct, reflection and voting. Also auditing an agent system somebody else built: tracks, evidence tiers and a prioritized plan instead of a score, plus a scanner. Carries Pi as a worked kernel implementation — SDK, RPC and extension seams, for embedding or extending a harness. Not for the loop's plumbing, its evals, or its protocols — those are siblings.


Install just this one

Every pack installs standalone. The whole family is one command.

any agent the skills CLI supports
$
npx skills add ssheleg/agent-stack
claude code, as a plugin
$
claude plugin marketplace add ssheleg/agent-stack && claude plugin install agent-stack@agent-stack

When the agent reaches for it

These are the rules the family writes into your agent's own instruction file — verbatim. Each one states the rule, the boundary in both directions, and the phrase that declines it.

/agent-stack — how an agent system is built, judged and metered — and the wire it speaks outward

  • when: the thing being built is an agent, or a server one connects to
  • decline it: no agent layer

If agent-stack is installed, building an agent SYSTEM goes through it — loops/work-graphs, system prompts/tools, workflow vs agent, evals, MCP/A2A, registry/gateway, and the wallet behind resold LLM access.

The boundary. NOT through it: one LLM call in a script, one prompt's wording, person-facing UI (super-ux, sheleg-design), charging money (sheleg-dev), agents editing THIS repo (agent-sync). sheleg-dev charges; agent-stack meters LLM use.

Refusal phrase: "no agent layer" or «без агентного слоя».

Among the routers: agent-sync coordinates file ownership; agent-stack owns the agent system being built.


The rest of the family

super-ux v0.59.1

what the interface must do

Scenario-driven UI development: a versioned design chain in docs/ux/ — the product vision, personas and jobs, user flows and the paid-acquisition funnel among them, a screens-and-states map with Figma frames…

7 skills/super-ux

task-pipeline v1.90.2

how a change reaches the repository

Full-cycle delivery orchestrator: an intake grill turns the request into a complete brief, then ten gated stages carry it from docs to acceptance, refusing to advance until each gate passes.

3 skills/task-pipeline

agent-sync v1.21.5

who is holding this file right now

Coordination plane for concurrent agents: leases with a TTL so two agents cannot claim the same work, race-free id reservation, a run journal and a generated board, over a pluggable knowledge cloud.

1 skill/agent-sync

make-skill v0.29.2

how the skill or plugin itself is built

A skill that builds skills: create, retrofit, audit and publish agent skills and Claude Code plugins

1 skill/make-skill

sheleg-design v1.64.0

how it looks and moves

The taste layer: cinematic scroll-driven landing pages (one scroll clock, motion that degrades to calm, WebGL particle formations), product-UI style packs with a ready token layer each, and the Figma border

1 skill · a catalogue of its own/sheleg-design

seo-aeo-audit v0.26.3

whether a machine will find it

Evidence-first website audit for search and answer engines: ten tracks from crawl access to AI citation mechanics, every finding backed by an observation and every recommendation tiered, ending in…

1 skillRead →

sheleg-dev v0.13.2

integrations: money in, tracking, errors, sign-in, speed

The integration layer a product reaches once it has users: Stripe subscription billing reconciled into your own database

7 skills/sheleg-dev

telegram-dev v0.2.2

which Telegram API each surface speaks, and what each one costs

Telegram, split by the API each surface actually speaks, in three skills. telegram-bots covers the official HTTP Bot API

3 skills/telegram-dev

xr-dev v0.3.2

how a Quest product moves from platform choice to launch and operation

Quest product delivery from platform choice to launch and operation, with stage owners, evidence gates and next actions.

7 skills · a catalogue of its own/xr-dev

web3d-dev v0.1.2

how realtime 3D on the web runs, ships its assets and moves

Realtime 3D on the web with three.js and React Three Fiber, in three skills split by the question. web3d-runtime owns how the scene runs

3 skills/web3d-dev