Create evaluation suite
POST/ai/studio/suites
Creates a new evaluation suite grouping a set of test cases and scoring criteria for benchmarking AI model outputs. Suites are used in the AI Studio to compare prompt template versions or to validate governance policy impact on response quality.
Request
Request body
Content type application/json (required).
| Field | Type | Required | Description |
|---|---|---|---|
name | string | Yes | Max length: 255 |
description | string | No | |
template_id | integer | No | |
scoring_criteria | array of object | No | |
scoring_criteria[].metric | string | Yes | Name of the evaluation metric (e.g. "accuracy", "conciseness"). |
scoring_criteria[].weight | number (float) | Yes | Weight of this metric in the aggregate score (weights must sum to 1.0). |
Example
{
"name": "Support Ticket Summary — Accuracy Suite",
"description": "Evaluates summary accuracy, hallucination rate, and SLA mention correctness.",
"template_id": 12,
"scoring_criteria": [
{
"metric": "accuracy",
"weight": 0.5
},
{
"metric": "hallucination_rate",
"weight": 0.3
},
{
"metric": "conciseness",
"weight": 0.2
}
]
}
Responses
201 Evaluation suite created
Content type application/json, object · AiEvaluationSuiteV1.
| Field | Type | Required | Description |
|---|---|---|---|
id | integer | No | |
name | string | No | Human-readable suite name. |
description | string, nullable | No | |
template_id | integer, nullable | No | ID of the prompt template this suite evaluates. |
scoring_criteria | array of object | No | List of scored metrics with weights (must sum to 1.0). |
scoring_criteria[].metric | string | No | |
scoring_criteria[].weight | number (float) | No | |
created_at | string (date-time) | No |
422 Validation failed
Example request
curl -X POST "https://{tenant}.faciotech.net/api/v1/ai/studio/suites" \
-H "Accept: application/json" \
-H "Content-Type: application/json" \
-d '{"name": "Support Ticket Summary — Accuracy Suite", "description": "Evaluates summary accuracy, hallucination rate, and SLA mention correctness.", "template_id": 12, "scoring_criteria": [{"metric": "accuracy", "weight": 0.5}, {"metric": "hallucination_rate", "weight": 0.3}, {"metric": "conciseness", "weight": 0.2}]}'