Unexpected behavior when using dist.all_reduce(x, op=dist.ReduceOp.SUM) #152300
Labels
module: c10d
Issues/PRs related to collective communications and process groups
oncall: distributed
Add this issue/PR to distributed oncall triage queue
triaged
This issue has been looked at a team member, and triaged and prioritized into an appropriate module
π Describe the bug
about 0.007% points didn't match.
Versions
python3.8.5
torch2.4.0
cc @H-Huang @awgu @wanchaol @fegin @fduwjj @wz337 @wconstab @d4l3k
The text was updated successfully, but these errors were encountered: