Grilled Cheese

ExploreLog inSign up

Explore

PostsPeople
LatestRanked
@tmlr-pub.bsky.socialOct 5, 2026, 4:30 PM

New #TMLR-Paper-with-Video:

A Multi-Fidelity Control Variate Approach for Policy Gradient Estimation

Xinjie Liu, Cyrus Neary, Kushagra Gupta et al.

https://tmlr.infinite-conf.org/paper_pages/zAo0L7Dcqt

#reinforcement #reinforce #trained

A Multi-Fidelity Control Variate Approach for Policy Gradient Estimation
@eralife.bsky.socialSep 22, 2026, 11:52 AM

2/By training ants with local vision and path context via #backpropagation and #REINFORCE, NAnts can collaboratively paint complex shapes like geckos, butterflies, and snails.

Even if partially erased or started anywhere on the grid, they dynamically repair the image! 🦎✨

Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT