[core] use swap_tensors in group offloading where possible
#3120
background
wait
wait-all
cancel
parallel
Loading