Agentβ™₯︎Age
Catalog

RAGScore

Official

by HZYAI Β· Python

Generate QA datasets & evaluate RAG systems with failure diagnosis. Any LLM.

io.github.HZYAI/ragscore generates QA datasets and evaluates RAG systems with failure diagnosis across any LLM. It provides tools for synthetic data generation, evaluation, and RAG pipeline analysis, suitable for local and Colab/Jupyter environments.

πŸ› οΈ Key Features

  • Generates QA datasets for RAG evaluation
  • Evaluates RAG systems with failure diagnosis
  • Works with any LLM and local deployments
  • Supports privacy-focused workflows
  • Includes tooling for dataset generation and fine-tuning considerations

πŸš€ Use Cases

  • RAG system evaluation and benchmarking
  • QA dataset creation for model testing
  • Fine-tuning data generation and evaluation
  • Privacy-conscious LLM experimentation
  • Local and Colab/Jupyter-based workflows

⚑ Developer Benefits

  • Wide compatibility with local LLMs and Colab/Jupyter
  • Clear evaluation metrics and failure analysis
  • Open-source licensing (Apache 2.0)
  • Easy integration with MCP (Model Context Protocol) pipelines
  • Active community topics: evaluation, privacy, synthetic data, and RAG

⚠️ Limitations

  • Documentation in readmeExcerpt is partial; full capabilities depend on repository contents
  • Requires compatible LLMs and environment setup for end-to-end runs
  • May necessitate data privacy considerations when generating synthetic data

Topics

evaluationollamaprivacyqa-generationragsynthetic-datadataset-generationfine-tuningllmlocal-llmcolabjupytermcprag-evaluationai-evaluationllmopsllm-as-a-judge