Flaglab/esnlir-test
Viewer • Updated • 80.2k • 20
Spanish NLI datasets with a causal 'reasoning' label, from ESNLIR. Used to evaluate open LLMs in github.com/Pacolas/NLI-via-LLM.
Note 1,971 pairs with human majority labels beside the connector-derived ones; the 999 disagreements behind the kappa = 0.33 result.