Skip to content

feat: add opt-in verbose_logs to surface GEval judge criteria, steps,… - #301

Open
x86girl wants to merge 5 commits into
lightspeed-core:mainfrom
x86girl:prgutier/g-eval/rubrics
Open

feat: add opt-in verbose_logs to surface GEval judge criteria, steps,…#301
x86girl wants to merge 5 commits into
lightspeed-core:mainfrom
x86girl:prgutier/g-eval/rubrics

Conversation

@x86girl

@x86girl x86girl commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Currently, lightspeed-evaluation does not report when creating rubrics using a Judge (LLM).
For correctness or consistency, for example, it asks the judge to define the steps to get this rubric, and lightspeed-evaluation is
not reporting anywhere how it has been done.
This PR create a new opt-in verbose_logs arg to surface GEval judge criteria, steps, and rubrics.

Summary by CodeRabbit

  • New Features

    • Added optional verbose mode for GEval metrics to capture detailed evaluation logs.
    • Exposed verbose logs in evaluation results and JSON/CSV exports.
    • Persisted verbose logs in SQL-stored results with automatic schema migration.
    • Supports both turn-level and conversation-level evaluations.
    • Combined logs from multiple judges when available.
  • Bug Fixes

    • Verbose logs now reset between evaluations and remain empty when disabled.

… and rubrics

GEval metrics ask the judge LLM to generate evaluation steps and rubrics
internally, but these were never exposed in the output — making it hard
to understand how a score was derived. This adds a `verbose: true` flag
to any GEval metric config that captures the judge's criteria, evaluation
steps, and rubric from DeepEval's verbose_logs and propagates them
through the full pipeline to the JSON/CSV output.

Changes:
- GEvalConfig: new `verbose` bool field (default False), parsed from metadata
- GEvalHandler: capture `metric.verbose_logs` at turn and conversation level
- DeepEvalMetrics: propagate last_verbose_logs from the GEval handler
- MetricResult: new `verbose_logs` optional field, inherited by EvaluationResult
- MetricsEvaluator: read handler verbose_logs after evaluation, attach to result
- Serializer: include verbose_logs in JSON output only when non-None
- constants: add verbose_logs to SUPPORTED_CSV_COLUMNS
- system.yaml: document the verbose option in metrics_metadata example

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: e9c2ffc6-f6be-4aef-8d7d-4f462e045292

📥 Commits

Reviewing files that changed from the base of the PR and between d146d2c and a0e39d4.

📒 Files selected for processing (4)
  • src/lightspeed_evaluation/core/storage/sql_storage.py
  • tests/unit/core/output/test_final_coverage.py
  • tests/unit/core/storage/test_sql_storage.py
  • tests/unit/pipeline/evaluation/test_judges.py
🚧 Files skipped from review as they are similar to previous changes (2)
  • tests/unit/core/output/test_final_coverage.py
  • src/lightspeed_evaluation/core/storage/sql_storage.py

Walkthrough

GEval now supports opt-in verbose logging from metadata configuration through handler capture, result propagation, JSON and CSV output, and SQL persistence. Tests cover defaults, turn and conversation capture, reset behavior, propagation, serialization, and schema migration.

Changes

GEval verbose log propagation

Layer / File(s) Summary
Verbose logging contracts
src/lightspeed_evaluation/core/models/llm.py, src/lightspeed_evaluation/core/models/data.py, src/lightspeed_evaluation/core/constants.py, config/system.yaml, tests/unit/core/models/*
GEval metadata accepts a verbose flag. Result models expose optional verbose_logs. CSV columns include verbose_logs. Configuration and model tests cover the new fields.
GEval log capture
src/lightspeed_evaluation/core/metrics/geval.py, src/lightspeed_evaluation/core/metrics/deepeval.py, tests/unit/core/metrics/test_geval.py
GEval handlers propagate the verbose setting, reset previous logs, capture turn- and conversation-level logs, and expose them through DeepEvalMetrics.
Result propagation and persistence
src/lightspeed_evaluation/pipeline/evaluation/judges.py, src/lightspeed_evaluation/core/output/serializers.py, src/lightspeed_evaluation/core/storage/sql_storage.py, tests/unit/pipeline/evaluation/test_evaluator.py, tests/unit/pipeline/evaluation/test_judges.py, tests/unit/core/output/test_final_coverage.py, tests/unit/core/storage/test_sql_storage.py
Judge scores and metric results collect verbose logs. JSON output always includes the field. CSV output preserves populated and absent values. SQL storage maps the field to a nullable text column and adds the column to legacy tables when needed. Tests validate propagation, serialization, and schema migration.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant GEvalConfig
  participant GEvalHandler
  participant DeepEvalMetrics
  participant JudgeOrchestrator
  participant result_to_json_dict
  participant EvaluationResultDB
  GEvalConfig->>GEvalHandler: provide verbose configuration
  GEvalHandler->>GEvalHandler: capture metric.verbose_logs
  DeepEvalMetrics->>GEvalHandler: read last_verbose_logs
  JudgeOrchestrator->>DeepEvalMetrics: evaluate judge
  JudgeOrchestrator->>JudgeOrchestrator: aggregate verbose_logs
  result_to_json_dict->>result_to_json_dict: include verbose_logs
  EvaluationResultDB->>EvaluationResultDB: store verbose_logs
Loading
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly identifies the main change: adding opt-in GEval verbose logs to expose judge details.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/lightspeed_evaluation/pipeline/evaluation/evaluator.py`:
- Around line 230-236: Update the GEval/DeepEval verbose-log propagation in the
evaluator to use logs from the handlers created by _create_handler_for_judge
rather than self.handlers[framework]. Carry the selected or aggregated panel
logs through JudgeOrchestrator and assign them to MetricResult for the
corresponding panel run, avoiding stale primary-handler state. Add a test
covering verbose-log propagation from a panel-created handler.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 51c7c3aa-fd25-42c7-8778-aa22a8d6ec65

📥 Commits

Reviewing files that changed from the base of the PR and between 590f807 and 84eff7c.

📒 Files selected for processing (13)
  • config/system.yaml
  • src/lightspeed_evaluation/core/constants.py
  • src/lightspeed_evaluation/core/metrics/deepeval.py
  • src/lightspeed_evaluation/core/metrics/geval.py
  • src/lightspeed_evaluation/core/models/data.py
  • src/lightspeed_evaluation/core/models/llm.py
  • src/lightspeed_evaluation/core/output/serializers.py
  • src/lightspeed_evaluation/pipeline/evaluation/evaluator.py
  • tests/unit/core/metrics/test_geval.py
  • tests/unit/core/models/test_data.py
  • tests/unit/core/models/test_system.py
  • tests/unit/core/output/test_final_coverage.py
  • tests/unit/pipeline/evaluation/test_evaluator.py

Comment thread src/lightspeed_evaluation/pipeline/evaluation/evaluator.py Outdated

@asamal4 asamal4 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks !!
Please address coderabbit's comment. and PTAL my comments..

Thread safety — shared mutable state between write and read
Judge panel path — verbose logs not captured
Only GEval — but verbose logs are available on all DeepEval metrics
Only file backends — missing from SQL, MLflow, Langfuse
Inconsistent JSON schema — conditionally omitted unlike other optional fields

Comment thread src/lightspeed_evaluation/core/metrics/geval.py
Comment thread src/lightspeed_evaluation/pipeline/evaluation/evaluator.py Outdated
Comment thread src/lightspeed_evaluation/core/models/data.py
Comment thread src/lightspeed_evaluation/core/output/serializers.py Outdated
Comment thread src/lightspeed_evaluation/core/metrics/deepeval.py
Comment thread config/system.yaml

@xmican10 xmican10 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! I noticed that the PR description mentions CSV export support but I don't see a test covering that path. Is it tested somewhere I missed?

x86girl and others added 2 commits July 30, 2026 12:20
…, DB column, JSON schema

- Move verbose_logs capture from evaluator (shared handler state) into
  JudgeOrchestrator._evaluate_single_judge (per-judge handler, no race)
- Propagate verbose_logs through JudgeScore → MetricResult (panel path works)
- Add verbose_logs column to EvaluationResultDB (SQL storage)
- Always include verbose_logs in JSON output (consistent with other optional fields)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The DB schema test was missing the verbose_logs column added in the
previous commit, causing CI to fail on all Python versions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/lightspeed_evaluation/core/storage/sql_storage.py`:
- Line 78: Update initialize() to apply an additive migration that adds the
nullable verbose_logs column to existing result databases before validating the
table schema. Preserve strict validation after migration and ensure the
migration is safe for databases where the column already exists; add coverage
for upgrading a pre-existing database without raising StorageError.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 18cb156d-ee34-40b3-b666-231bee5684d6

📥 Commits

Reviewing files that changed from the base of the PR and between 84eff7c and b56812d.

📒 Files selected for processing (6)
  • src/lightspeed_evaluation/core/models/data.py
  • src/lightspeed_evaluation/core/output/serializers.py
  • src/lightspeed_evaluation/core/storage/sql_storage.py
  • src/lightspeed_evaluation/pipeline/evaluation/judges.py
  • tests/unit/core/output/test_final_coverage.py
  • tests/unit/core/storage/test_sql_storage.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • src/lightspeed_evaluation/core/models/data.py

Comment thread src/lightspeed_evaluation/core/storage/sql_storage.py
Existing databases created before this PR are missing the verbose_logs
column and would fail initialize() schema validation.  Add a migration
step that runs ALTER TABLE ADD COLUMN for any missing nullable columns
before validation, and refresh the inspector cache afterward.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/lightspeed_evaluation/core/storage/sql_storage.py`:
- Around line 173-193: Update the schema initialization flow around
`_migrate_missing_columns` and `_validate_evaluation_results_schema` to
preflight missing non-nullable columns before applying any ALTER TABLE
statements. Reject non-result or structurally incompatible `evaluation_results`
tables first, and only run `_migrate_missing_columns` when all required ORM
columns are present; preserve nullable-column migration for compatible legacy
tables.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 9315fb3c-ce26-470b-a913-f4709e5134c5

📥 Commits

Reviewing files that changed from the base of the PR and between b56812d and d146d2c.

📒 Files selected for processing (1)
  • src/lightspeed_evaluation/core/storage/sql_storage.py

Comment thread src/lightspeed_evaluation/core/storage/sql_storage.py
@asamal4

asamal4 commented Jul 30, 2026

Copy link
Copy Markdown
Collaborator

@x86girl is this PR ready ? There is a open coderabbit comment.

…ges tests

Addresses final CodeRabbit review: reject incompatible schemas before
ALTER TABLE.  Adds CSV export test (xmican10 feedback) and verbose_logs
propagation tests for single/multi-judge paths.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@x86girl

x86girl commented Aug 12, 2026

Copy link
Copy Markdown
Contributor Author

@xmican10 Good catch. CSV export test for verbose_logs has been added in a0e39d4 (TestCsvExportVerboseLogs in tests/unit/core/output/test_final_coverage.py). It verifies the column appears in CSV output with the correct value when set, and as empty when not set.

@asamal4 All review comments have been addressed:

  • Thread safety: verbose logs captured per-judge in _evaluate_single_judge, no shared mutable state
  • Judge panel path: logs propagated through JudgeScore into MetricResult
  • SQL storage: verbose_logs column added with additive migration + preflight check for incompatible schemas
  • JSON schema: verbose_logs always included as null when not set
  • CSV test: added
  • GEval-only scope: acknowledged, follow-up PR for all DeepEval metrics

All CI checks pass. PR is ready for re-review.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants