> ## Documentation Index
> Fetch the complete documentation index at: https://docs.4minds.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Evaluate your Model Performance

> Evaluations provide comprehensive model performance reports with key metrics to help you make data-driven decisions about model quality. The 4MINDS platform automatically generates evaluation reports that assess your model's accuracy, response quality, and overall effectiveness. Use Evaluations feature to identify areas for improvement and validate that your model meets production requirements.

To evaluate your model, navigate to the '**Evaluations**' tab and click '**Create Evaluation**'.

<img alt="Screen Shot2025 10 31at1 34 59PM Pn" lightAlt="Screen Shot2025 10 31at1 34 59PM Pn" darkAlt="Screen Shot2025 10 31at1 34 59PM Pn" src="https://mintcdn.com/4minds-e9525117/SKhMgATBbfIefDq5/images/Screenshot-2026-06-10-at-5.47.15-PM.png?fit=max&auto=format&n=SKhMgATBbfIefDq5&q=85&s=db67df3693ad35160727dd98917f69ac" className="mx-auto dark:hidden" width="2008" height="412" data-path="images/Screenshot-2026-06-10-at-5.47.15-PM.png" />

<img alt="Screen Shot2025 10 31at1 34 59PM Pn" lightAlt="Screen Shot2025 10 31at1 34 59PM Pn" darkAlt="Screen Shot2025 10 31at1 34 59PM Pn" src="https://mintcdn.com/4minds-e9525117/XFeMmj16ehmQCYCk/images/Screenshot-2026-06-10-at-4.59.38-PM.png?fit=max&auto=format&n=XFeMmj16ehmQCYCk&q=85&s=315cd4b131640872073c83a562c3b701" className="mx-auto hidden dark:block" width="2038" height="460" data-path="images/Screenshot-2026-06-10-at-4.59.38-PM.png" />

Select your evaluation method from the available options:

* **RAGAS Benchmark** (***Currently Available***): A standardized test that measures model performance on specific tasks using predefined metrics and datasets.
* **Model as Judge** (***Currently Available***): Automatically compares your customized model against a base foundation model, with ChatGPT acting as an AI judge to evaluate responses side-by-side and determine which performs better.
* **Model Comparison** (***Coming Soon***)**:** Side-by-side evaluation of multiple models using standardized tests to compare performance, accuracy, and response quality across specific tasks.
* **Human Evaluation** (***Coming Soon***): Manual assessment by human reviewers to evaluate response quality, relevance, and appropriateness based on subjective criteria.

## RAGAS Benchmark evaluation

**Step 1: Choose 'RAGAS Benchmark' evaluation method**

<img alt="Screen Shot2025 11 25at2 28 12PM Pn" title="Screen Shot2025 11 25at2 28 12PM Pn" style={{ width:"58%" }} lightAlt="Screen Shot2025 11 25at2 28 12PM Pn" darkAlt="Screen Shot2025 11 25at2 28 12PM Pn" src="https://mintcdn.com/4minds-e9525117/SKhMgATBbfIefDq5/images/Screenshot-2026-06-10-at-5.48.06-PM.png?fit=max&auto=format&n=SKhMgATBbfIefDq5&q=85&s=54b54aee39134aad1d6f8a1846abd0d7" className="mx-auto dark:hidden" width="558" height="800" data-path="images/Screenshot-2026-06-10-at-5.48.06-PM.png" />

<img alt="Screen Shot2025 11 25at2 28 12PM Pn" title="Screen Shot2025 11 25at2 28 12PM Pn" style={{ width:"58%" }} lightAlt="Screen Shot2025 11 25at2 28 12PM Pn" darkAlt="Screen Shot2025 11 25at2 28 12PM Pn" src="https://mintcdn.com/4minds-e9525117/XFeMmj16ehmQCYCk/images/Screenshot-2026-06-10-at-5.04.15-PM.png?fit=max&auto=format&n=XFeMmj16ehmQCYCk&q=85&s=e020a8ad8e2e01142d431823ddd90bdc" className="mx-auto hidden dark:block" width="574" height="820" data-path="images/Screenshot-2026-06-10-at-5.04.15-PM.png" />

Click '**Next**' to continue.

**Step 2: Select your base model**

<Note>
  **Direct base-model selection has been deprecated.** Models deployed directly on the 4MINDS platform run on `gpt-oss-120b`, which is also the default comparison baseline used in evaluations. To compare against other foundation models, connect them through an external provider integration like [Amazon Bedrock](/bedrock), Google Vertex AI, [Amazon SageMaker](/integrations#amazon-sagemaker), or [Microsoft Foundry](/microsoft-foundry).
</Note>

<img alt="Screen Shot2025 11 25at2 37 42PM Pn" title="Screen Shot2025 11 25at2 37 42PM Pn" style={{ width:"55%" }} lightAlt="Screen Shot2025 11 25at2 37 42PM Pn" darkAlt="Screen Shot2025 11 25at2 37 42PM Pn" src="https://mintcdn.com/4minds-e9525117/SKhMgATBbfIefDq5/images/Screenshot-2026-06-10-at-5.48.57-PM.png?fit=max&auto=format&n=SKhMgATBbfIefDq5&q=85&s=a783351b60760c107bb4181b21ad0a3f" className="mx-auto dark:hidden" width="576" height="450" data-path="images/Screenshot-2026-06-10-at-5.48.57-PM.png" />

<img alt="Screen Shot2025 11 25at2 37 42PM Pn" title="Screen Shot2025 11 25at2 37 42PM Pn" style={{ width:"55%" }} lightAlt="Screen Shot2025 11 25at2 37 42PM Pn" darkAlt="Screen Shot2025 11 25at2 37 42PM Pn" src="https://mintcdn.com/4minds-e9525117/XFeMmj16ehmQCYCk/images/Screenshot-2026-06-10-at-5.06.32-PM.png?fit=max&auto=format&n=XFeMmj16ehmQCYCk&q=85&s=2bb4730af57e7b2ddcc13d002d56a065" className="mx-auto hidden dark:block" width="566" height="584" data-path="images/Screenshot-2026-06-10-at-5.06.32-PM.png" />

Click '**Next**' to continue.

**Step 3: Review and confirm the details**

<Frame>
  <img src="https://mintcdn.com/4minds-e9525117/SKhMgATBbfIefDq5/images/Screenshot-2026-06-10-at-5.50.00-PM.png?fit=max&auto=format&n=SKhMgATBbfIefDq5&q=85&s=5306aa50578a94a5469de33256cc9354" alt="Screenshot 2026 06 10 At 5 07 29 PM" title="Screenshot 2026 06 10 At 5 07 29 PM" lightAlt="Screenshot 2026 06 10 At 5 07 29 PM" darkAlt="Screenshot 2026 06 10 At 5 07 29 PM" className="mx-auto dark:hidden" width="568" height="1012" data-path="images/Screenshot-2026-06-10-at-5.50.00-PM.png" />

  <img src="https://mintcdn.com/4minds-e9525117/XFeMmj16ehmQCYCk/images/Screenshot-2026-06-10-at-5.07.29-PM.png?fit=max&auto=format&n=XFeMmj16ehmQCYCk&q=85&s=4a356510806be569aa58f2bccf9902fe" alt="Screenshot 2026 06 10 At 5 07 29 PM" title="Screenshot 2026 06 10 At 5 07 29 PM" lightAlt="Screenshot 2026 06 10 At 5 07 29 PM" darkAlt="Screenshot 2026 06 10 At 5 07 29 PM" className="mx-auto hidden dark:block" width="572" height="1060" data-path="images/Screenshot-2026-06-10-at-5.07.29-PM.png" />
</Frame>

Review and confirm the details, then click '**Start Evaluation**' to begin.

Monitor the evaluation status in the Evaluation Dashboard. Once complete, the status will update to "**Completed**" and the evaluation report will open in a separate window.

<img alt="Screen Shot2025 10 31at1 51 40PM Pn" title="Screenshot2026 02 17at7 55 39PM" style={{ width:"61%" }} className="mx-auto dark:hidden" lightAlt="Screen Shot2025 10 31at1 51 40PM Pn" darkAlt="Screen Shot2025 10 31at1 51 40PM Pn" src="https://mintcdn.com/4minds-e9525117/SKhMgATBbfIefDq5/images/Screenshot-2026-06-10-at-5.51.09-PM.png?fit=max&auto=format&n=SKhMgATBbfIefDq5&q=85&s=6f04a173bf60c50901ca59cc80f9885a" width="1904" height="1566" data-path="images/Screenshot-2026-06-10-at-5.51.09-PM.png" />

<img alt="Screen Shot2025 10 31at1 51 40PM Pn" title="Screenshot2026 02 17at7 55 39PM" style={{ width:"61%" }} className="mx-auto hidden dark:block" lightAlt="Screen Shot2025 10 31at1 51 40PM Pn" darkAlt="Screen Shot2025 10 31at1 51 40PM Pn" src="https://mintcdn.com/4minds-e9525117/XFeMmj16ehmQCYCk/images/Screenshot-2026-06-10-at-5.10.52-PM.png?fit=max&auto=format&n=XFeMmj16ehmQCYCk&q=85&s=bb745a043093a2d255489b423502a213" width="1902" height="1572" data-path="images/Screenshot-2026-06-10-at-5.10.52-PM.png" />

## Model as Judge evaluation

Model as Judge automatically compares your customized model against a base foundation model. ChatGPT evaluates responses side-by-side to determine which performs better.

### Setting up a Model as Judge evaluation

**Step 1: Choose 'Model as Judge' evaluation method**

<img alt="Screen Shot2025 11 25at2 26 40PM Pn" title="Screen Shot2025 11 25at2 26 40PM Pn" style={{ width:"46%" }} lightAlt="Screen Shot2025 11 25at2 26 40PM Pn" darkAlt="Screen Shot2025 11 25at2 26 40PM Pn" src="https://mintcdn.com/4minds-e9525117/SKhMgATBbfIefDq5/images/Screenshot-2026-06-10-at-5.51.58-PM.png?fit=max&auto=format&n=SKhMgATBbfIefDq5&q=85&s=34df25c13c99829b398d24d1d342a76d" className="mx-auto dark:hidden" width="566" height="812" data-path="images/Screenshot-2026-06-10-at-5.51.58-PM.png" />

<img alt="Screen Shot2025 11 25at2 26 40PM Pn" title="Screen Shot2025 11 25at2 26 40PM Pn" style={{ width:"46%" }} lightAlt="Screen Shot2025 11 25at2 26 40PM Pn" darkAlt="Screen Shot2025 11 25at2 26 40PM Pn" src="https://mintcdn.com/4minds-e9525117/XFeMmj16ehmQCYCk/images/Screenshot-2026-06-10-at-5.14.08-PM.png?fit=max&auto=format&n=XFeMmj16ehmQCYCk&q=85&s=b6617a59a38f92ea1b7a10a92c269c8e" className="mx-auto hidden dark:block" width="580" height="808" data-path="images/Screenshot-2026-06-10-at-5.14.08-PM.png" />

Click '**Next**' to continue.

**Step 2: Select your trained model**

<img alt="Screen Shot2025 11 24at3 20 17PM Pn" title="Screen Shot2025 11 24at3 20 17PM Pn" style={{ width:"38%" }} lightAlt="Screen Shot2025 11 24at3 20 17PM Pn" darkAlt="Screen Shot2025 11 24at3 20 17PM Pn" src="https://mintcdn.com/4minds-e9525117/SKhMgATBbfIefDq5/images/Screenshot-2026-06-10-at-5.52.42-PM.png?fit=max&auto=format&n=SKhMgATBbfIefDq5&q=85&s=593eb59c37e52b8340f3cc8c21714117" className="mx-auto dark:hidden" width="570" height="752" data-path="images/Screenshot-2026-06-10-at-5.52.42-PM.png" />

<img alt="Screen Shot2025 11 24at3 20 17PM Pn" title="Screen Shot2025 11 24at3 20 17PM Pn" style={{ width:"38%" }} lightAlt="Screen Shot2025 11 24at3 20 17PM Pn" darkAlt="Screen Shot2025 11 24at3 20 17PM Pn" src="https://mintcdn.com/4minds-e9525117/XFeMmj16ehmQCYCk/images/Screenshot-2026-06-10-at-5.15.13-PM.png?fit=max&auto=format&n=XFeMmj16ehmQCYCk&q=85&s=2e50f4e00344fc6d29b60eeafabbd395" className="mx-auto hidden dark:block" width="572" height="754" data-path="images/Screenshot-2026-06-10-at-5.15.13-PM.png" />

Select the customized model you want to evaluate, review its description, and click '**Next**' to continue.

**Step 3: Review and confirm**

<img alt="Screen Shot2025 11 24at3 20 49PM Pn" title="Screen Shot2025 11 24at3 20 49PM Pn" style={{ width:"44%" }} lightAlt="Screen Shot2025 11 24at3 20 49PM Pn" darkAlt="Screen Shot2025 11 24at3 20 49PM Pn" src="https://mintcdn.com/4minds-e9525117/SKhMgATBbfIefDq5/images/Screenshot-2026-06-10-at-5.53.14-PM.png?fit=max&auto=format&n=SKhMgATBbfIefDq5&q=85&s=426552f109a64ed255bb2828e8af9d1d" className="mx-auto dark:hidden" width="556" height="842" data-path="images/Screenshot-2026-06-10-at-5.53.14-PM.png" />

<img alt="Screen Shot2025 11 24at3 20 49PM Pn" title="Screen Shot2025 11 24at3 20 49PM Pn" style={{ width:"44%" }} lightAlt="Screen Shot2025 11 24at3 20 49PM Pn" darkAlt="Screen Shot2025 11 24at3 20 49PM Pn" src="https://mintcdn.com/4minds-e9525117/XFeMmj16ehmQCYCk/images/Screenshot-2026-06-10-at-5.17.24-PM.png?fit=max&auto=format&n=XFeMmj16ehmQCYCk&q=85&s=daf27d8ef71e5d9867264148cd09e255" className="mx-auto hidden dark:block" width="562" height="860" data-path="images/Screenshot-2026-06-10-at-5.17.24-PM.png" />

Verify your settings:

* **Evaluation Name**: Auto-generated name
* **Your Trained Model**: Your customized model (with RAG)
* **Base Model**: Foundation model for comparison (without RAG)
* **Evaluation Type**: Model as Judge

Click **Start Comparison** to begin.

### Using the evaluation interface

<img alt="Screen Shot2025 11 24at3 22 00PM Pn" title="Screen Shot2025 11 24at3 22 00PM Pn" style={{ width:"71%" }} lightAlt="Screen Shot2025 11 24at3 22 00PM Pn" darkAlt="Screen Shot2025 11 24at3 22 00PM Pn" src="https://mintcdn.com/4minds-e9525117/SKhMgATBbfIefDq5/images/Screenshot-2026-06-10-at-5.53.56-PM.png?fit=max&auto=format&n=SKhMgATBbfIefDq5&q=85&s=db230bcf793ebf88a3cd1ac3a8627f3e" className="mx-auto dark:hidden" width="2102" height="1560" data-path="images/Screenshot-2026-06-10-at-5.53.56-PM.png" />

<img alt="Screen Shot2025 11 24at3 22 00PM Pn" title="Screen Shot2025 11 24at3 22 00PM Pn" style={{ width:"71%" }} lightAlt="Screen Shot2025 11 24at3 22 00PM Pn" darkAlt="Screen Shot2025 11 24at3 22 00PM Pn" src="https://mintcdn.com/4minds-e9525117/XFeMmj16ehmQCYCk/images/Screenshot-2026-06-10-at-5.21.12-PM.png?fit=max&auto=format&n=XFeMmj16ehmQCYCk&q=85&s=97b8f18592be94dd62dfafe484065c5d" className="mx-auto hidden dark:block" width="2288" height="1532" data-path="images/Screenshot-2026-06-10-at-5.21.12-PM.png" />

The interface displays a side-by-side chat comparison:

* **Left panel**: Your trained model (with your data)
* **Right panel**: Base model (without your data)

**To test your models:**

1. Type your question in the input box
2. Send to both models simultaneously
3. Review responses in real-time
4. Scroll down to view automated evaluation results

### Understanding evaluation results

Each evaluation summary includes:

**Winner declaration**\
Shows which model provided the better response

**Factual grounding analysis**

* Response A (RAG): How well your model uses training data
* Response B (Base): Evaluation of the unenhanced model

**Key differences**\
Highlights why one response outperformed the other

**Winner rationale**\
Detailed explanation of the judge's decision

### Evaluation criteria

The AI judge evaluates responses based on:

* Factual accuracy from source material
* Proper use of grounding and training data
* Relevance to the question
* Completeness and clarity

<Note>
  Grounded responses using your training data consistently outperform speculative answers.
</Note>

### Interpreting your results

**Your model wins**\
Your customization is working effectively. Training data is being used properly, and the model is ready for this use case.

**Base model wins**\
A knowledge gap has been identified. Add more training data on this topic and continue refinement.

**Mixed results**\
Partial success indicates you should add data for questions where your model underperformed and continue testing.

### Best practices

**Testing strategy:**

* Ask 10-15 diverse questions minimum
* Test scenarios where your data should provide an advantage
* Include difficult and edge cases
* Review "Past Results" to track improvement over time

**After evaluation:**

1. Identify patterns in wins and losses
2. Add training data to address knowledge gaps
3. Re-test to verify improvements
4. Iterate continuously

### Tips for success

* Grounded responses always outperform speculation
* Losses reveal where to add more training data
* Test regularly as you add new content
* Use realistic queries your actual users would ask
