Dashboard
Signal #141856NEUTRAL

Show HN: LitigationBench. A Litigation Task-Based AI Benchmark

80

Adding to the sea of AI benchmarks, but with a focus on legal and litigation-based tasks and tests. E.g., hallucinating cases, misreading precedent, drafting, AI writing tics, etc.Developed as part of a litigation platform I've started, but not sharing this benchmark to promote that. Just thought the findings were interesting--namely, Anthropic's models (except for Haiku) producing zero hallucinated cases.Also, if anyone would like to see additional tests / benchmarks or models tested, let me know and I'll incorporate them into v3 if it makes sense. Comments URL: https://news.ycombinator.com/item?id=49015617 Points: 1 # Comments: 0

HackerNews Show AIabout 4 hours ago
Read Full Article

Explore with AI-Powered Tools

View All Signals

Explore more AI intelligence

Want to discover more AI signals like this?

Explore Steek
Show HN: LitigationBench. A Litigation Task-Based AI Benchmark | Steek AI Signal | Steek