Use case · decision ranking
Evaluation and quality gates
Run repeatable evaluations or checks before software or AI changes advance.
5 reviewed matches, ranked by fit and deterministic project health.
#1
Editorialwebdriverio/webdriverio: Next-gen browser and mobile automation test framework for Node.js.
Editorial94% fitThe reviewed testing category fits repeatable evaluation and quality gates around software or AI behavior.Compare this repository → ★ 9.8KTypeScriptMITbrowser-automationclitesting
#2
Editorialconfident-ai/deepeval: The LLM Evaluation Framework.
Editorial94% fitThe reviewed testing category fits repeatable evaluation and quality gates around software or AI behavior.Compare this repository → ★ 17.8KPythonApache-2.0clitesting
#3
Editoriallaravel/dusk: Laravel Dusk provides simple end-to-end testing and browser automation.
Editorial94% fitThe reviewed testing category fits repeatable evaluation and quality gates around software or AI behavior.Compare this repository → ★ 1.9KPHPMITbrowser-automationclitesting
#4
Editorialopenai/evals: Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Editorial94% fitThe reviewed testing category fits repeatable evaluation and quality gates around software or AI behavior.Compare this repository → ★ 19.2KPythonNOASSERTIONclitesting
#5
EditorialServiceNow/BrowserGym: 🌎💪 BrowserGym, a Gym environment for web task automation.
Editorial94% fitThe reviewed testing category fits repeatable evaluation and quality gates around software or AI behavior.Compare this repository → ★ 1.3KPythonNOASSERTIONbrowser-automationclitesting