benchmarking-clinical-ner
maziyarpanahi/openmed
This skill provides an in-depth, entity-level scorecard for clinical or biomedical Named Entity Recognition (NER) models. It calculates precision, recall, and F1 scores (both strict and partial match modes) using a user-supplied gold corpus. Key features include per-label error breakdown, a confusion matrix, and detailed False Negative/False Positive analysis, essential for debugging model performance in highly sensitive medical domains.