Skip to content
View sixijsu77-hub's full-sized avatar

Block or report sixijsu77-hub

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Popular repositories Loading

  1. mimir-korean-retrieval mimir-korean-retrieval Public

    Pre-registered, reproducible measurement of BM25 / dense / hybrid retrieval on public Korean benchmarks (MTEB-ko). The harness is validated against published baselines before any new number is repo…

    Python

  2. themis-judge-reliability themis-judge-reliability Public

    Do LLM judges and reward models still agree with themselves when the same question is asked with the predicate inverted? Pre-registered measurement on RewardBench 2, run on one 24 GB GPU with open …

    Python

  3. reward-bench reward-bench Public

    Forked from allenai/reward-bench

    RewardBench: the first evaluation tool for reward models.

    Python

  4. mteb mteb Public

    Forked from embeddings-benchmark/mteb

    MTEB: State-of-the-art evaluation of embeddings across languages and modalities

    Python