技能 人工智能 Hugging Face 模型评估管理

Hugging Face 模型评估管理

v20260927
hugging-face-evaluation
此技能用于在 Hugging Face 模型卡片中添加和管理结构化评估结果。支持从 README 内容中提取评估表格,通过 Artificial Analysis API 导入基准分数,并使用 vLLM 或 lighteval 后端运行自定义模型评估。确保与 model-index 元数据格式兼容,适用于模型发布准备。
获取技能
90 次下载
概览

Overview

This skill provides tools to add structured evaluation results to Hugging Face model cards. It supports multiple methods for adding evaluation data:

  • Extracting existing evaluation tables from README content
  • Importing benchmark scores from Artificial Analysis
  • Running custom model evaluations with vLLM or accelerate backends (lighteval/inspect-ai)

Detailed Guide

Read the detailed guide before executing this skill. It retains the complete procedure and reference material. Treat its safety, prerequisites, and validation requirements as mandatory. For focused work, load the relevant sections; for end-to-end work, read the guide completely.

When to Use

  • You need to add structured evaluation results to a Hugging Face model card.
  • You want to import benchmark data or run custom evaluations with vLLM, lighteval, or inspect-ai.
  • You are preparing leaderboard-compatible model-index metadata for a model release.

Limitations

  • Use this skill only when the task clearly matches the scope described above.
  • Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
  • Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
信息
Category 人工智能
Name hugging-face-evaluation
版本 v20260927
大小 8.23KB
更新时间 2026-09-28
语言