Files
my-vault/01_Projects/Personal-Tech/LLM_Evaluation/04-Reference-Archive/evaluation-guidebook/00-Overview.md
T
windyboy f375cd6134 add evaluation-guidebook notes: HF guidebook knowledge (2024 GitHub + 2025 Space) into LLM_Evaluation
- Add 9 Chinese distilled notes under 04-Reference-Archive/evaluation-guidebook/
  covering automated benchmarks, human evaluation, LLM-as-a-judge,
  troubleshooting, general knowledge, yearly dives, resources, and the
  2025 HF Space edition (FineWeb eval-selection methodology, MCF/CF/FG,
  sampling metrics, Math-Verify, statistical validity & cost)
- Update zone README and Material-List with entry links
2026-08-21 16:19:12 +08:00

5.6 KiB
Raw Blame History

type, tags, status, created, source
type tags status created source
reference
llm-evaluation
evaluation-guidebook
overview
active 2026-08-21 https://github.com/huggingface/evaluation-guidebook

HuggingFace LLM Evaluation Guidebook 总览

外部权威参考的来源说明:本子目录的内容是对 HuggingFace Evaluation Guidebook(旧版,GitHub 仓库)与 OpenEvals 新版(2025HuggingFace Space)的中文提炼笔记,按主题拆成独立页面,供本专区在设计与执行评测时查阅。

这是什么

The LLM Evaluation Guidebook 是 HuggingFace 团队(主要作者 Clémentine FourrierOpen LLM Leaderboard 与 lighteval 的设计者)编写的 LLM 评测实战指南。它回答一个核心问题:

如何确保一个 LLM 在你自己的具体任务上表现良好?

内容覆盖:评测模型的不同方式、如何设计自己的评测、以及从实际评测工程中沉淀的经验教训(Tips and Tricks)。

两个版本

旧版(GitHub 仓库) 新版(HF Space2025-12
地址 github.com/huggingface/evaluation-guidebook huggingface.co/spaces/OpenEvals/evaluation-guidebook
形态 章节式 Markdown 指南 交互式"科研论文"Astro + MDX,含可交互图表)
作者 Clémentine Fourrier Fourrier、Thibaud Frere、Guilherme Penedo、Thomas Wolf
副标题 "基于 3 年评测 15000 个模型的经验,你想知道的关于 LLM 评测的一切"
状态 已停止维护(README 声明) 当前维护版本(2025-12-03 发布,CC BY 4.0
本专区笔记 01-Automatic-Benchmarks ~ 07-Resources 08-2025-Edition(新版独有内容提炼)

⚠️ 维护状态:旧仓库已声明不再维护(截至 2025 年 12 月),最新版本迁移到 HuggingFace Space。本目录 0107 篇整理的是旧版 GitHub 仓库内容;08-2025-Edition 提炼新版(2025)相对旧版新增/变化的内容。两份内容大部分重叠,建议以新版为主、旧版为补充。

指南结构(旧版仓库目录)

章节 内容 本专区对应笔记
Automatic benchmarks 自动化基准评测:basics、设计自己的自动评测、常用评测数据集、Tips 01-Automatic-Benchmarks
Human evaluation 人工评测:basics、如何使用标注者、Tips 02-Human-Evaluation
LLM-as-a-judge 模型作评委:basics、选择 judge LLM、设计评测 prompt、评估你的 evaluator、奖励模型 03-LLM-as-a-Judge
Troubleshooting 指南中最实操的部分:推理排错、LaTeX 数学解析、可复现性 04-Troubleshooting
General knowledge LLM 基础:模型推理与评测、tokenization 05-General-Knowledge
Yearly dives 2023/2024/2025 年度深度文章 06-Yearly-Dives
Resources 评测与 NLP 推荐链接清单 07-Resources

新版结构(2025 Space,按渲染顺序)

章节 内容 备注
Intro 评测视角:model builder vs model user、智能定义的困境 新版新增
Model inference and evaluation Tokenization、推理、MCF/CF/FG 三种任务形式、calibration 比旧版更细
2025 evaluations 2025 分能力评测全景(推理/知识/数学/代码/长上下文/指令遵循/工具调用/游戏化) 新版核心,替代旧版 Yearly Dives 2025
Troubleshooting reproducibility 可复现性排错(代码库/种子/指标名/normalization/prompt/参数) 旧版有对应章节
Picking good automatic evaluations for pretraining FineWeb 团队预训练评测选型方法论(185 任务、SNR、单调性、排序一致性) 全新内容
Designing your automatic evaluation 设计自动评测全流程(数据集/提示词/推理方式/评分/自由文本评分/约束输出/统计有效性/成本) 内含 Using human annotators
Conclusion 五条核心建议

注:仓库中 some-evaluation-datasets.mdx 存在但新版页面未直接渲染(正文链接回旧版 GitHub 的同类页面);using-human-annotators.mdx 通过 import 嵌入 designing 章节。

建议阅读方式

原作者的建议(旧版):

  • 初学者:从每个章节的 Basics 部分开始,需要 LLM 基础补课时读 General knowledge
  • 进阶用户:直接看每个章节的 Tips and TricksTroubleshooting 章节。
  • 回访用户:每年一篇的 Yearly dives(每年的主题深潜)。

文内标记 的链接是作者特别推荐阅读的资源。

与本专区的关系

  • 本专区(LLM_Evaluation)是面向个人学习与工程能力养成的中文路线(概念 → 为什么 → 工作表 → 项目 → 完整路线图)。
  • 本子目录是外部权威知识参考:需要深入某个评测主题(如设计自动评测、写 judge prompt、排查复现问题)时,从这里查阅提炼后的要点,并按需回到原文细读。
  • 两者互补:专区路线图负责"做什么、按什么顺序做"guidebook 笔记负责"具体怎么做、有哪些坑"。

引用

旧版(GitHub):

@misc{fourrier2024evaluation,
  author = {Clémentine Fourrier and The Hugging Face Community},
  title = {LLM Evaluation Guidebook},
  year = {2024},
  journal = {GitHub repository},
  url = {https://github.com/huggingface/evaluation-guidebook}
}

许可证:旧版 CC BY-NC-SA 4.0;新版 CC BY 4.0。