Skip to content

Evaluate multiclass logloss on the CUDA device - #7472

Closed
yananlong wants to merge 1 commit into
lightgbm-org:mainfrom
yananlong:codex/cuda-multiclass-metric
Closed

yananlong wants to merge 1 commit into
lightgbm-org:mainfrom
yananlong:codex/cuda-multiclass-metric

Conversation

@yananlong

Copy link
Copy Markdown

Summary

CUDA multi_logloss currently falls back to the CPU, which requires copying the multiclass score matrix back from the device at evaluation time. Add a CUDA reduction for softmax and one-vs-all objectives, including sample weights, and compare the reported metric with predictions in the CUDA test suite.

Validation

  • git diff --check and Python syntax checks pass.
  • An earlier combined HIP/gfx1100 branch passed seven CUDA/HIP Python regression tests. The new weighted and one-vs-all cases have not yet run on this isolated current-main branch, so this PR is a draft.
  • A prior single paired timing run showed 61.355 s for the correctness baseline and 56.917 s for the branch containing both this metric and the tree-transfer candidate. That combined result does not attribute the change to this PR; no standalone speedup claim is made here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants