Compare frameworks
Ragas vs DSPy vs DeepEval
Every figure below carries the source that reported it and the moment it was observed. Values that failed validation are withheld rather than shown, and a dash means no source we track carries that field for that framework.
| Field | Ragas | DSPy | DeepEval |
|---|---|---|---|
| Stars | 15,862Synced from sourceCurrent5h ago | 38,388 (highest of the compared values)Synced from sourceCurrent5h ago | 18,476Synced from sourceCurrent5h ago |
| Last commit | 2026-02-24Synced from sourceCurrent5h ago | 2026-09-27Synced from sourceCurrent5h ago | 2026-09-25Synced from sourceCurrent5h ago |
| Language | pythonSeeded, unreviewedwritten 1mo ago | pythonSeeded, unreviewedwritten 1mo ago | pythonSeeded, unreviewedwritten 1mo ago |
| Licence | Apache-2.0Synced from sourceCurrent5h ago | MITSynced from sourceCurrent5h ago | Apache-2.0Synced from sourceCurrent5h ago |
| Archived | noSynced from sourceCurrent5h ago | noSynced from sourceCurrent5h ago | noSynced from sourceCurrent5h ago |
| Category | evaluationSeeded, unreviewedwritten 1mo ago | optimisationSeeded, unreviewedwritten 1mo ago | evaluationSeeded, unreviewedwritten 1mo ago |
| Contributors | 246Synced from sourceCurrent5h ago | 461Synced from sourceCurrent5h ago | 337Synced from sourceCurrent5h ago |
| Forks | 1728Synced from sourceCurrent5h ago | 3372Synced from sourceCurrent5h ago | 1989Synced from sourceCurrent5h ago |
| Maintainer | Vibrant LabsSeeded, unreviewedwritten 1mo ago | Stanford NLPSeeded, unreviewedwritten 1mo ago | Confident AISeeded, unreviewedwritten 1mo ago |
| Open issues | 617Synced from sourceCurrent5h ago | 739Synced from sourceCurrent5h ago | 676Synced from sourceCurrent5h ago |
| PyPI package | ragasSynced from sourceCurrent5h ago | dspy-aiSynced from sourceCurrent5h ago | deepevalSynced from sourceCurrent5h ago |
| PyPI released | 2026-01-13T17:47:59.200116ZSynced from sourceCurrent5h ago | 2026-09-25T04:04:30.339197ZSynced from sourceCurrent5h ago | 2026-09-24T09:31:55.310624ZSynced from sourceCurrent5h ago |
| Requires Python | >=3.9Synced from sourceCurrent5h ago | >=3.9Synced from sourceCurrent5h ago | <4.0,>=3.9Synced from sourceCurrent5h ago |
| PyPI version | 0.4.3Synced from sourceCurrent5h ago | 3.4.0Synced from sourceCurrent5h ago | 4.2.6Synced from sourceCurrent5h ago |
| Repository | https://github.com/vibrantlabsai/ragasSynced from sourceCurrent5h ago | https://github.com/stanfordnlp/dspySynced from sourceCurrent5h ago | https://github.com/confident-ai/deepevalSynced from sourceCurrent5h ago |
Sources: GitHub REST API, npm registry downloads, PyPI registry metadata, pypistats.org downloads. Most recent observation 5h ago. Where a row is marked, ▲ is the highest of the values shown and ▼ the lowest — arithmetic on the figures above, not a ranking or a recommendation. Rows where neither extreme is meaningful are left unmarked rather than given a direction they do not have.
Change the comparison
Remove one, or open the directory to pick a different set. Up to 3 at a time.