skills/smithery.ai/langfuse-experiment-runner

langfuse-experiment-runner

Installation
SKILL.md

Langfuse Experiment Runner

Run experiments on datasets with evaluators stored in Langfuse or custom scripts. Analyze results and compare runs.

Key Feature: Judge prompts can be stored in Langfuse for versioning and reuse across experiments.

Score Scale Policy

  • Canonical score scale is 0-1.
  • If judge outputs 0-10, runner normalizes to 0-1 before aggregation and threshold checks.
  • Reports may display both representations (e.g., 0.82 (8.2/10)), but gating logic uses canonical 0-1.

When to Use

  • Running experiments on Langfuse datasets
  • Evaluating prompt/model changes against test sets
  • Comparing multiple experiment runs
  • Analyzing score distributions and failures
  • Building regression test workflows
Installs
1
First Seen
Mar 30, 2026
langfuse-experiment-runner from smithery.ai