Example: biology
Evaluating Large Language Models Trained on Code

Evaluating Large Language Models Trained on Code

Back to document page

human evaluators. To accurately benchmark our model, we create a dataset of 164 original programming problems with unit tests. These problems assess language compre-hension, algorithms, and simple mathematics, with some comparable to simple software interview questions. We release this data along with an evaluation framework at

  Human, Evaluating

Download Evaluating Large Language Models Trained on Code


Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Related search queries