일단 Evaluation를 정의하면, Model 객체나 LLM 애플리케이션 로직이 포함된 사용자 정의 함수에 대해 실행할 수 있습니다. .evaluate()에 대한 각 호출은 evaluation run을 트리거합니다. Evaluation 객체를 청사진으로, 각 실행을 해당 설정에서 애플리케이션이 어떻게 수행되는지에 대한 측정으로 생각하세요.평가를 시작하려면 다음 단계를 완료하세요:
객체 Evaluation를 생성하는 것은 평가 구성을 설정하는 첫 번째 단계입니다. Evaluation는 예제 데이터, 점수 매기기 로직 및 선택적 전처리로 구성됩니다. 나중에 이를 사용하여 하나 이상의 평가를 실행하게 됩니다.Weave는 각 예제를 가져와 애플리케이션을 통해 전달한 다음 여러 사용자 정의 점수 매기기 함수로 출력 점수를 매깁니다. 이를 통해 애플리케이션의 성능을 볼 수 있고, 개별 출력 및 점수를 자세히 살펴볼 수 있는 풍부한 UI를 갖게 됩니다.
먼저, Dataset 객체 또는 평가할 예제 모음이 있는 딕셔너리 목록을 정의합니다. 이러한 예제는 종종 테스트하려는 실패 사례로, 테스트 주도 개발(TDD)의 단위 테스트와 유사합니다.다음 예제는 딕셔너리 목록으로 정의된 데이터셋을 보여줍니다:
examples = [ {"question": "What is the capital of France?", "expected": "Paris"}, {"question": "Who wrote 'To Kill a Mockingbird'?", "expected": "Harper Lee"}, {"question": "What is the square root of 64?", "expected": "8"},]
그런 다음, 하나 이상의 점수 매기기 함수를 생성합니다. 이는 Dataset의 각 예제에 점수를 매기는 데 사용됩니다. 각 점수 매기기 함수는 반드시 output를 가져야 하며, 점수가 포함된 딕셔너리를 반환해야 합니다. 선택적으로 예제에서 다른 입력을 포함할 수 있습니다.점수 매기기 함수는 output 키워드 인수를 가져야 하지만, 다른 인수는 사용자 정의이며 데이터셋 예제에서 가져옵니다. 인수 이름을 기반으로 한 딕셔너리 키를 사용하여 필요한 키만 가져옵니다.
점수 매기기가 output 인수를 예상하지만 받지 못하는 경우, 레거시 model_output 키를 사용하고 있는지 확인하세요. 이를 수정하려면 점수 매기기 함수를 업데이트하여 output을 키워드 인수로 사용하세요.
다음 예제 점수 매기기 함수 match_score1는 expected 값을 examples 딕셔너리에서 점수 매기기에 사용합니다.
import weave# Collect your examplesexamples = [ {"question": "What is the capital of France?", "expected": "Paris"}, {"question": "Who wrote 'To Kill a Mockingbird'?", "expected": "Harper Lee"}, {"question": "What is the square root of 64?", "expected": "8"},]# Define any custom scoring function@weave.op()def match_score1(expected: str, output: dict) -> dict: # Here is where you'd define the logic to score the model output return {'match': expected == output['generated_text']}
일부 애플리케이션에서는 사용자 정의 Scorer 클래스를 생성하고자 합니다 - 예를 들어 특정 매개변수(예: 채팅 모델, 프롬프트), 각 행의 특정 점수 매기기, 그리고 집계 점수의 특정 계산이 포함된 표준화된 LLMJudge 클래스를 생성해야 하는 경우가 있습니다.다음에서 Scorer 클래스 정의에 관한 튜토리얼을 참조하세요: Model-Based Evaluation of RAG applications에서 더 많은 정보를 확인할 수 있습니다.
다음을 평가하려면 Model, evaluate를 Evaluation를 사용하여 호출하세요. Models는 실험하고 weave에 캡처하려는 매개변수가 있을 때 사용됩니다.
from weave import Model, Evaluationimport asyncioclass MyModel(Model): prompt: str @weave.op() def predict(self, question: str): # here's where you would add your LLM call and return the output return {'generated_text': 'Hello, ' + self.prompt}model = MyModel(prompt='World')evaluation = Evaluation( dataset=examples, scorers=[match_score1])weave.init('intro-example') # begin tracking results with weaveasyncio.run(evaluation.evaluate(model))
이는 각 예제에 대해 predict를 실행하고 각 점수 매기기 함수로 출력에 점수를 매깁니다.
@weave.opdef function_to_evaluate(question: str): # here's where you would add your LLM call and return the output return {'generated_text': 'some response'}asyncio.run(evaluation.evaluate(function_to_evaluate))
다음 코드 샘플은 처음부터 끝까지 완전한 평가 실행을 보여줍니다. examples 딕셔너리는 match_score1와 match_score2 점수 매기기 함수에서 MyModel의 값이 주어진 prompt를 평가하는 데 사용되며, 사용자 정의 함수 function_to_evaluate도 평가합니다. Model와 함수에 대한 평가 실행은 asyncio.run(evaluation.evaluate()를 통해 호출됩니다.
from weave import Evaluation, Modelimport weaveimport asyncioweave.init('intro-example')examples = [ {"question": "What is the capital of France?", "expected": "Paris"}, {"question": "Who wrote 'To Kill a Mockingbird'?", "expected": "Harper Lee"}, {"question": "What is the square root of 64?", "expected": "8"},]@weave.op()def match_score1(expected: str, output: dict) -> dict: return {'match': expected == output['generated_text']}@weave.op()def match_score2(expected: dict, output: dict) -> dict: return {'match': expected == output['generated_text']}class MyModel(Model): prompt: str @weave.op() def predict(self, question: str): # here's where you would add your LLM call and return the output return {'generated_text': 'Hello, ' + question + self.prompt}model = MyModel(prompt='World')evaluation = Evaluation(dataset=examples, scorers=[match_score1, match_score2])asyncio.run(evaluation.evaluate(model))@weave.op()def function_to_evaluate(question: str): # here's where you would add your LLM call and return the output return {'generated_text': 'some response' + question}asyncio.run(evaluation.evaluate(function_to_evaluate("What is the capitol of France?")))
우리는 타사 서비스 및 라이브러리와의 통합을 지속적으로 개선하고 있습니다.더 원활한 통합을 구축하는 동안, preprocess_model_input를 Weave 평가에서 HuggingFace Datasets를 사용하기 위한 임시 해결책으로 사용할 수 있습니다.현재 접근 방식은 Using HuggingFace datasets in evaluations cookbook를 참조하세요.