Classification evaluation evidence
Evaluate up to 1,000 caller-supplied single-label expected/predicted pairs: confusion matrix, per-label precision/recall/F1/support, macro/weighted/micro summaries, balanced accuracy and mismatch evidence. Undefined per-label values use zero with explicit flags; labels are not verified ground truth.
data-assurance · Operation ID: classification-evaluate
Choose this operation when
- calculate confusion matrix precision recall f1 from evaluation labels
- score agent routing classification predictions against supplied labels
- inspect macro micro weighted metrics and error samples
Outside this profile
- Train or execute a model or infer factual ground truth
- Multilabel classification probability calibration or deployment certification
Exact release references
Static JSON contract · Markdown reference · Fixed example response
Complete input schema
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"samples": {
"maxItems": 1000,
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string",
"minLength": 1,
"maxLength": 80
},
"expected": {
"type": "string",
"minLength": 1,
"maxLength": 80
},
"predicted": {
"type": "string",
"minLength": 1,
"maxLength": 80
}
},
"required": [
"id",
"expected",
"predicted"
],
"additionalProperties": false
}
},
"labels": {
"default": [],
"maxItems": 50,
"type": "array",
"items": {
"type": "string",
"minLength": 1,
"maxLength": 80
}
}
},
"required": [
"samples"
],
"additionalProperties": false
}
Complete output-envelope schema
{
"type": "object",
"required": [
"operation",
"version",
"result",
"provenance"
],
"properties": {
"operation": {
"const": "classification-evaluate",
"type": "string"
},
"version": {
"const": "0.29.0",
"type": "string"
},
"result": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"profile": {
"type": "string",
"const": "single-label-evaluation-v1"
},
"sampleCount": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"labels": {
"type": "array",
"items": {
"type": "string"
}
},
"confusionMatrix": {
"type": "array",
"items": {
"type": "array",
"items": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
}
}
},
"perLabel": {
"type": "array",
"items": {
"type": "object",
"properties": {
"precision": {
"type": "number"
},
"recall": {
"type": "number"
},
"f1": {
"type": "number"
},
"label": {
"type": "string"
},
"truePositive": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"falsePositive": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"falseNegative": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"support": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"predictedCount": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"precisionDefined": {
"type": "boolean"
},
"recallDefined": {
"type": "boolean"
},
"f1Defined": {
"type": "boolean"
}
},
"required": [
"precision",
"recall",
"f1",
"label",
"truePositive",
"falsePositive",
"falseNegative",
"support",
"predictedCount",
"precisionDefined",
"recallDefined",
"f1Defined"
],
"additionalProperties": false
}
},
"accuracy": {
"type": [
"number",
"null"
]
},
"balancedAccuracy": {
"type": [
"number",
"null"
]
},
"macro": {
"anyOf": [
{
"type": "object",
"properties": {
"precision": {
"type": "number"
},
"recall": {
"type": "number"
},
"f1": {
"type": "number"
}
},
"required": [
"precision",
"recall",
"f1"
],
"additionalProperties": false
},
{
"type": "null"
}
]
},
"weighted": {
"anyOf": [
{
"type": "object",
"properties": {
"precision": {
"type": "number"
},
"recall": {
"type": "number"
},
"f1": {
"type": "number"
}
},
"required": [
"precision",
"recall",
"f1"
],
"additionalProperties": false
},
{
"type": "null"
}
]
},
"micro": {
"anyOf": [
{
"type": "object",
"properties": {
"precision": {
"type": "number"
},
"recall": {
"type": "number"
},
"f1": {
"type": "number"
}
},
"required": [
"precision",
"recall",
"f1"
],
"additionalProperties": false
},
{
"type": "null"
}
]
},
"mismatchCount": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"mismatches": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string"
},
"expected": {
"type": "string"
},
"predicted": {
"type": "string"
}
},
"required": [
"id",
"expected",
"predicted"
],
"additionalProperties": false
}
},
"mismatchesTruncated": {
"type": "boolean"
},
"conventions": {
"type": "array",
"items": {
"type": "string"
}
}
},
"required": [
"profile",
"sampleCount",
"labels",
"confusionMatrix",
"perLabel",
"accuracy",
"balancedAccuracy",
"macro",
"weighted",
"micro",
"mismatchCount",
"mismatches",
"mismatchesTruncated",
"conventions"
],
"additionalProperties": false
},
"provenance": {
"type": "object",
"required": [
"inputSha256",
"outputSha256",
"deterministic",
"externalRequests"
],
"properties": {
"inputSha256": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"outputSha256": {
"type": "string",
"pattern": "^[a-f0-9]{64}$"
},
"deterministic": {
"const": true
},
"externalRequests": {
"const": 0
}
}
}
},
"additionalProperties": false
}
Fixed example
One accepted fixed example, not a custom-input trial. No operation runs when this static page is requested.
Example input
{
"samples": [
{
"id": "e1",
"expected": "answer",
"predicted": "answer"
},
{
"id": "e2",
"expected": "abstain",
"predicted": "answer"
},
{
"id": "e3",
"expected": "abstain",
"predicted": "abstain"
}
],
"labels": [
"answer",
"abstain"
]
}
Example response
{
"operation": "classification-evaluate",
"version": "0.29.0",
"result": {
"profile": "single-label-evaluation-v1",
"sampleCount": 3,
"labels": [
"abstain",
"answer"
],
"confusionMatrix": [
[
1,
1
],
[
0,
1
]
],
"perLabel": [
{
"precision": 1,
"recall": 0.5,
"f1": 0.6666666666666666,
"label": "abstain",
"truePositive": 1,
"falsePositive": 0,
"falseNegative": 1,
"support": 2,
"predictedCount": 1,
"precisionDefined": true,
"recallDefined": true,
"f1Defined": true
},
{
"precision": 0.5,
"recall": 1,
"f1": 0.6666666666666666,
"label": "answer",
"truePositive": 1,
"falsePositive": 1,
"falseNegative": 0,
"support": 1,
"predictedCount": 2,
"precisionDefined": true,
"recallDefined": true,
"f1Defined": true
}
],
"accuracy": 0.6666666666666666,
"balancedAccuracy": 0.75,
"macro": {
"precision": 0.75,
"recall": 0.75,
"f1": 0.6666666666666666
},
"weighted": {
"precision": 0.8333333333333334,
"recall": 0.6666666666666666,
"f1": 0.6666666666666666
},
"micro": {
"precision": 0.6666666666666666,
"recall": 0.6666666666666666,
"f1": 0.6666666666666666
},
"mismatchCount": 1,
"mismatches": [
{
"id": "e2",
"expected": "abstain",
"predicted": "answer"
}
],
"mismatchesTruncated": false,
"conventions": [
"Confusion matrix rows are expected labels and columns are predicted labels; labels use code-unit lexical order.",
"Undefined per-label precision, recall or F1 is 0 with explicit defined flags. Macro includes all declared and observed labels; weighted uses expected support; balanced accuracy includes only supported labels.",
"Supplied expectations are not verified. No training, probability calibration, multilabel matching or inference about deployment quality."
]
},
"provenance": {
"inputSha256": "94e0ebdf6418835b5e24379c85e0dcf742165925f5cc787097b124279db65c32",
"outputSha256": "96304c42a7fa68f16782023dec0075af56722a1a75d391c578fcedb286492dc3",
"deterministic": true,
"externalRequests": 0
}
}
Bounds and precision
JavaScript IEEE-754 numbers; use strings for large integer IDs/exact decimals where the schema accepts strings. No lossless numeric parsing.
{
"global": {
"requestBytes": 131072,
"responseBytes": 524288,
"jsonDepth": 32,
"jsonNodes": 20000,
"requestsPerMinute": 60,
"paidAttemptsPerMinute": 20,
"idempotencyHours": 24
},
"operation": {
"inputBytes": 100000,
"outputBytes": 400000,
"jsonNodes": 12000,
"depth": 20,
"evidence": 200,
"keyWorkBytes": 3000000,
"matchingSteps": 2000000,
"samples": 1000,
"labels": 50
}
}
Complete schemas, descriptions and cross-field validation may impose additional limits.
Proposed price and protocol definitions
{
"unit": "one successful operation call",
"proposedNominalUsd": "0.01",
"sixDecimalTokenBaseUnits": "10000",
"subscription": false,
"includesPayerWalletOrNetworkFees": false,
"liveQuoteVerified": false,
"condition": "Actual SDK challenge is authoritative only within the caller's explicit authorization; configured six-decimal token peg is an operator assertion, not a conversion guarantee."
}
Protocol definitions: x402, mpp. MPP uses Tempo charge. Paid MCP execution is unsupported. All runtime readiness is not evaluated in this build.
API path templates, not endpoints on this documentation host
{
"x402": "/v1/x402/classification-evaluate",
"mpp": "/v1/mpp/classification-evaluate"
}
Required headers
{
"Content-Type": "application/json",
"Idempotency-Key": "random 16–128 character operation identifier"
}
Actual SDK challenge amount, asset, network, recipient and wallet costs must pass independent authorization. Preserve identical key, body, protocol and credential on retries; on PAYMENT_UNCERTAIN stop and reconcile.
Execution profile and provider conditions
{
"deterministic": true,
"externalRequests": 0,
"maxExternalRequests": 0,
"resultSnapshotPersisted": false,
"fixedExampleIsIllustrativeSnapshot": false,
"requiresPayment": true,
"supportsMcpExecution": false
}
Deterministic supplied-input operation with no external requests or stored request/result bodies. Payment infrastructure retains payment metadata and hashes.
Failure handling
- HTTP 400: Malformed JSON, missing/invalid idempotency key, or payment identifier mismatch Correct the request before payment
- HTTP 402: Payment challenge or rejected payment Use official protocol SDK; inspect payment outcome before another payment
- HTTP 409: Idempotency conflict, duplicate proof, or PAYMENT_UNCERTAIN Keep original key, body, and proof; reconcile uncertainty with operator; never blindly repay
- HTTP 413: Input or generated output too large Reduce input; no payment attempted for validation failure
- HTTP 415: Unsupported media type or compression Send uncompressed application/json
- HTTP 422: Schema or service-specific semantic validation failure Correct input using returned error code; no payment attempted
- HTTP 429: Request/payment-attempt rate exceeded Wait for rate limit window; preserve existing payment identity
- HTTP 503: Payment configuration/provider/state unavailable, or live DNS preparation failed before settlement Check readiness; DNS preparation failures may retry the identical key/body/credential only; uncertainty requires reconciliation
Declared requirements
Before any paid call, refresh the live operation contract and POST the complete bounded budgeted plan to the separate API's /preflight. Unknown requirements block selection; compatible preflight is not permission to spend.