Weave의 로컬 스코러는 최소한의 지연 시간으로 사용자의 기기에서 로컬로 실행되는 소형 언어 모델 모음입니다. 이러한 모델은 AI 시스템의 입력, 컨텍스트 및 출력의 안전성과 품질을 평가합니다.이러한 모델 중 일부는 Weights & Biases에 의해 미세 조정되었으며, 다른 일부는 커뮤니티에서 훈련한 최첨단 오픈 소스 모델입니다. 훈련 및 평가에는 Weights & Biases(W&B) Reports가 사용되었습니다. 전체 세부 정보는 다음에서 확인할 수 있습니다 W&B Reports 목록.모델 가중치는 W&B Artifacts에서 공개적으로 사용 가능하며, 스코러 클래스를 인스턴스화할 때 자동으로 다운로드됩니다. 직접 다운로드하려면 다음에서 아티팩트 경로를 찾을 수 있습니다: weave.scorers.default_models이러한 스코러가 반환하는 객체에는 passed 입력 텍스트가 안전하거나 고품질인지 나타내는 boolean 속성과 metadata 모델의 원시 점수와 같은 더 자세한 정보를 포함하는 속성이 있습니다.
로컬 스코러는 CPU와 GPU에서 실행할 수 있지만, 최상의 성능을 위해 GPU를 사용하세요.
import weavefrom weave.scorers import WeaveBiasScorerV1bias_scorer = WeaveBiasScorerV1()result = bias_scorer.score(output="Martian men are terrible at cleaning")print(f"The text is biased: {not result.passed}")print(result)
import weavefrom weave.scorers import WeaveToxicityScorerV1toxicity_scorer = WeaveToxicityScorerV1()result = toxicity_scorer.score(output="people from the south pole of Mars are the worst")print(f"Input is toxic: {not result.passed}")print(result)
import weavefrom weave.scorers import WeaveHallucinationScorerV1hallucination_scorer = WeaveHallucinationScorerV1()result = hallucination_scorer.score( query="What is the capital of Antarctica?", context="People in Antarctica love the penguins.", output="While Antarctica is known for its sea life, penguins aren't liked there.")print(f"Output is hallucinated: {not result.passed}")print(result)
import weavefrom weave.scorers import WeaveContextRelevanceScorerV1context_relevance_scorer = WeaveContextRelevanceScorerV1()result = context_relevance_scorer.score( query="What is the capital of Antarctica?", output="The Antarctic has the happiest penguins." # context is passed to the output parameter)print(f"Output is relevant: {result.passed}")print(result)
import weavefrom weave.scorers import WeaveCoherenceScorerV1coherence_scorer = WeaveCoherenceScorerV1()result = coherence_scorer.score( query="What is the capital of Antarctica?", output="but why not monkey up day")print(f"Output is coherent: {result.passed}")print(result)
이 평가기는 입력 텍스트가 유창한지—즉, 자연스러운 인간 언어와 유사하게 읽고 이해하기 쉬운지 평가합니다. 문법, 구문 및 전반적인 가독성을 평가합니다.The WeaveFluencyScorerV1 uses a fine-tuned ModernBERT-base model from AnswerDotAI. For more information, see the WeaveFluencyScorerV1 W&B Report.
The WeaveTrustScorerV1 is a composite scorer for RAG systems that evaluates the trustworthiness of model outputs by grouping other scorers into two categories: Critical and Advisory. Based on the composite score, it returns a trust level:
high: 문제가 감지되지 않음
medium: Advisory 문제만 감지됨
low: Critical 문제가 감지되거나 입력이 비어 있음
Critical 평가기에서 실패한 모든 입력은 low trust level로 결과가 나옵니다. Advisory 평가기에서 실패하면 medium.
import weavefrom weave.scorers import WeaveTrustScorerV1trust_scorer = WeaveTrustScorerV1()def print_trust_scorer_result(result): print() print(f"Output is trustworthy: {result.passed}") print(f"Trust level: {result.metadata['trust_level']}") if not result.passed: print("Triggered scorers:") for scorer_name, scorer_data in result.metadata['raw_outputs'].items(): if not scorer_data.passed: print(f" - {scorer_name} did not pass") print() print(f"WeaveToxicityScorerV1 scores: {result.metadata['scores']['WeaveToxicityScorerV1']}") print(f"WeaveHallucinationScorerV1 scores: {result.metadata['scores']['WeaveHallucinationScorerV1']}") print(f"WeaveContextRelevanceScorerV1 score: {result.metadata['scores']['WeaveContextRelevanceScorerV1']}") print(f"WeaveCoherenceScorerV1 score: {result.metadata['scores']['WeaveCoherenceScorerV1']}") print(f"WeaveFluencyScorerV1: {result.metadata['scores']['WeaveFluencyScorerV1']}") print()result = trust_scorer.score( query="What is the capital of Antarctica?", context="People in Antarctica love the penguins.", output="The cat stretched lazily in the warm sunlight.")print_trust_scorer_result(result)print(result)
import weavefrom weave.scorers import PresidioScorerpresidio_scorer = PresidioScorer()result = presidio_scorer.score( output="Mary Jane is a software engineer at XYZ company and her email is mary.jane@xyz.com.")print(f"Output contains PII: {not result.passed}")print(result)
Weave 로컬 평가기는 아직 TypeScript에서 사용할 수 없습니다. 계속 지켜봐 주세요!TypeScript에서 Weave 평가기를 사용하려면 function-based scorers를 참조하세요.
이 페이지가 도움이 되었나요?
Assistant
Responses are generated using AI and may contain mistakes.