An approval without a test is a signature, and a signature is theatre. IJk records which agents run in your company, keeps what they answered, and puts two versions side by side with the numbers attached. So there is evidence under your decision.
A platform I built myself called a language model directly from nine separate files, each with its own model and its own temperature. Nobody had ever declared those nine as agents. They just did their work, day after day, and no one could say what they cost or how often they were wrong. That is the normal state in most companies working with AI.
Every agent gets an owner, a purpose in plain language and a risk class. An agent running without ever having been declared shows up at the top of the register as a finding.
Every run becomes a tree of steps with input, answer, cost and duration. Sensitive input disappears after its retention window while the step stays, so a run remains auditable after the content is gone.
You build a test set from real runs, change a prompt or a model, and see per case what got better, what got worse and what broke. With cost and duration next to it.
I put IJk on my own platform first and had one task scored by four models. Nine tasks from Flemish SME practice, with the category a person assigns to them.
The most expensive model is the only one that got everything right, at eighty-eight times the cost of the cheapest. Whether that is worth it depends on what rides on the decision. Without this table I would have argued about it. Now I chose.
The first real experiment scored zero percent. The model was fine: the agent produced strict JSON in production and prose under measurement, because the answer schema was not sent along. You never find that difference by looking at one answer. A second measurement exposed a model failing nine times out of nine on a parameter it no longer accepts.
IJk works with any model and any provider. Anything that speaks OpenTelemetry can report in.
IJk is tied to no platform. Anything that speaks OpenTelemetry can report in, and for code you write yourself it is one function around your model call.
Want to see what your agents do?
I set IJk up for you, usually within a day, and then we look at your first register together.
The first agent IJk flagged as undeclared was one of mine.