Thanks to visit codestin.com
Credit goes to github.com

Skip to content
#

llm-evaluation

Here are 200 public repositories matching this topic...

langfuse

๐Ÿชข Open source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more. ๐ŸŠYC W23

  • Updated May 6, 2025
  • TypeScript

Test your prompts, agents, and RAGs. Red teaming, pentesting, and vulnerability scanning for LLMs. Compare performance of GPT, Claude, Gemini, Llama, and more. Simple declarative configs with command line and CI/CD integration.

  • Updated May 6, 2025
  • TypeScript

Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. ไธ€ไธช็”ฑๅทฅๅ…ทใ€ๅŸบๅ‡†/ๆ•ฐๆฎใ€ๆผ”็คบใ€ๆŽ’่กŒๆฆœๅ’Œๅคงๆจกๅž‹็ญ‰็ป„ๆˆ็š„็ฒพ้€‰ๅˆ—่กจ๏ผŒไธป่ฆ้ขๅ‘ๅŸบ็ก€ๅคงๆจกๅž‹่ฏ„ๆต‹๏ผŒๆ—จๅœจๆŽขๆฑ‚็”ŸๆˆๅผAI็š„ๆŠ€ๆœฏ่พน็•Œ.

  • Updated Oct 25, 2024

Improve this page

Add a description, image, and links to the llm-evaluation topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the llm-evaluation topic, visit your repo's landing page and select "manage topics."

Learn more