Skip to content

Repeat fake_atom_weight by multiplicity in diffusion loss - #251

Open
fnachon wants to merge 9 commits into
HannesStark:mainfrom
fnachon:fix/diffusion-loss-fake-atom-weight
Open

Repeat fake_atom_weight by multiplicity in diffusion loss#251
fnachon wants to merge 9 commits into
HannesStark:mainfrom
fnachon:fix/diffusion-loss-fake-atom-weight

Conversation

@fnachon

@fnachon fnachon commented Jul 10, 2026

Copy link
Copy Markdown

fake_atom_weight is derived from feats["fake_atom_mask"], which is at the pre-multiplicity batch size, but it's multiplied against tensors (denoised_atom_coords, resolved_atom_mask, etc.) that are already repeated by multiplicity earlier in the same function. The mismatch is silent whenever multiplicity == 1 since a singleton dimension broadcasts fine against any size, which is presumably why it went unnoticed. Fixed by repeat_interleave-ing it the same way resolved_atom_mask and atom_type already are just above.

fnachon and others added 9 commits January 10, 2026 15:54
Changes made to run without errors on the Mac MPS device: torch.autocast, number of devices and workers to use on M1-5 chips, workaround for CUDA-specific code, handling of float64 incompatibilities for MPS.
Replace hardcoded torch.autocast("cuda") with device-agnostic
device_type=tensor.device.type in confidence_utils, inverse_fold,
and writer modules introduced in the upstream merge.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Python pickle does not preserve RDKit atom-level SetProp values. When
PyTorch DataLoader spawns worker processes (default num_workers=1 on
macOS), self.canonicals is pickled and all atom 'name' properties are
lost, causing KeyError in process_atom_features.

Fix: load all required molecules directly from the moldir zip inside
each get_sample() / get_feat() call instead of using the pickled
self.canonicals. The moldir zip handle is cached per-process by
_get_zipfile(), so there is no repeated I/O overhead.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…ders

- Disable pin_memory on MPS (unsupported, causes UserWarning)
- Enable persistent_workers when num_workers > 0 (avoids repeated
  worker init overhead and the PL suggestion warning)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
fake_atom_weight is derived from feats["fake_atom_mask"], which is at
the pre-multiplicity batch size, but it's multiplied against tensors
(denoised_atom_coords, resolved_atom_mask, ...) that are repeated by
`multiplicity`. The mismatch was silent whenever multiplicity == 1
because a singleton dimension broadcasts fine.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant