Repository navigation
README and docs homepage: results first, and the end-to-end chain from LLM to failure handling - #39
Merged
Merged
Conversation
Lead with what the toolkit does for AI-assisted curation and the measured results with six models; show one real run from LLM tool calls through structured output, evaluation and feedback, and a failure-handling table mapped to the code; state the limits and the negative result. Add a results figure rendered from the committed summaries (plot.py).
The site's home page now leads with the same results, the end-to-end chain and the failure-handling table, with site-relative links; the single-cell case page joins the benchmarks section.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rewrites the homepage README for readers assessing the project. It now leads with:
The figure
docs/assets/ai_validation_results.svgis rendered from the committed summaries byevaluation/llm_benchmark/plot.py. Its palette was checked with the dataviz validator, and every segment is labelled.Every number and example is taken from committed results:
results/agent-pilot;results/celltype-test.Links stay absolute, so the README also renders on PyPI.
Also syncs the docs site homepage (
docs/index.md) with the same content and site-relative links, and publishes the single-cell case as a benchmarks page (mkdocs_hooks.py,mkdocs.yml).mkdocs build --strictpasses.