Tag
1 verified claim carrying this tag. Each cites primary evidence; the source count is shown on every record.
HumanEval benchmark introduced in paper: Evaluating Large Language Models Trained on Code (Chen et al., 2021).
71ec42731d2c9e0c · 2 sources · 100% confidence