Skip to content

Create evaluation suite

POST/ai/studio/suites

Creates a new evaluation suite grouping a set of test cases and scoring criteria for benchmarking AI model outputs. Suites are used in the AI Studio to compare prompt template versions or to validate governance policy impact on response quality.

Request

Request body

Content type application/json (required).

FieldTypeRequiredDescription
namestringYes

Max length: 255

descriptionstringNo
template_idintegerNo
scoring_criteriaarray of objectNo
scoring_criteria[].metricstringYes

Name of the evaluation metric (e.g. "accuracy", "conciseness").

scoring_criteria[].weightnumber (float)Yes

Weight of this metric in the aggregate score (weights must sum to 1.0).

Example

{
  "name": "Support Ticket Summary — Accuracy Suite",
  "description": "Evaluates summary accuracy, hallucination rate, and SLA mention correctness.",
  "template_id": 12,
  "scoring_criteria": [
    {
      "metric": "accuracy",
      "weight": 0.5
    },
    {
      "metric": "hallucination_rate",
      "weight": 0.3
    },
    {
      "metric": "conciseness",
      "weight": 0.2
    }
  ]
}

Responses

201 Evaluation suite created

Content type application/json, object · AiEvaluationSuiteV1.

FieldTypeRequiredDescription
idintegerNo
namestringNo

Human-readable suite name.

descriptionstring, nullableNo
template_idinteger, nullableNo

ID of the prompt template this suite evaluates.

scoring_criteriaarray of objectNo

List of scored metrics with weights (must sum to 1.0).

scoring_criteria[].metricstringNo
scoring_criteria[].weightnumber (float)No
created_atstring (date-time)No

422 Validation failed

Example request

curl -X POST "https://{tenant}.faciotech.net/api/v1/ai/studio/suites" \
  -H "Accept: application/json" \
  -H "Content-Type: application/json" \
  -d '{"name": "Support Ticket Summary — Accuracy Suite", "description": "Evaluates summary accuracy, hallucination rate, and SLA mention correctness.", "template_id": 12, "scoring_criteria": [{"metric": "accuracy", "weight": 0.5}, {"metric": "hallucination_rate", "weight": 0.3}, {"metric": "conciseness", "weight": 0.2}]}'
Loading