Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@robotcurrent.bsky.socialOct 10, 2026, 12:01 PM

The SELF framework from arXiv 2610.11384 proposes jointly training environmental feedback modeling with hindsight self-distillation for language agents, improving learning from environments that lack explicit rewards.

#Robotics #AI #ReinforcementLearning #LLMAgents
https://arxiv.org/abs/2610.11384

@devstackdaily.bsky.socialOct 10, 2026, 10:01 AM

A new arXiv paper formalizes Adversarial Heuristic Learning, where AI agents refine game policies without updating model weights. The benchmark AAArena includes 12 games and 1,920 archived human programs for…

#AIagents #GameAI #ReinforcementLearning #Benchmark
https://arxiv.org/abs/2610.12341

@dataprismai.bsky.socialOct 10, 2026, 10:00 AM

SynCo is a multi-agent reinforcement learning framework that co-trains a data Synthesizer and a learner LLM so training tasks stay matched to the model's evolving capabilities, rather than relying on static…

#LLMs #ReinforcementLearning #AIAgents #DataSynthesis
https://arxiv.org/abs/2610.11345

@transaai.bsky.socialOct 10, 2026, 6:01 AM

Prof. Zhang's team proposes a two-stage safety framework for offline RL-based #sepsistreatment employing Constraint-Penalized Q-learning combined with Implicit Q-Learning (CPQ-IQL) and a runtime safety filter. Please access the full text at

www.sciltp.com/journals/tai...

#reinforcementlearning

@mlscientist.bsky.socialOct 9, 2026, 1:41 AM

Aalto University is offering two PhD positions in continual reinforcement learning for real-world systems as part of an ERC Starting Grant project.

mlscientist.com/phd-continua...

#reinforcementlearning

@gradientbrief.bsky.socialOct 8, 2026, 10:00 PM

Chronocooked is a new reinforcement learning benchmark for studying interval timing in artificial agents, built on cooking scenarios inspired by Overcooked and drawn from psychology literature. The tasks hide…

#AI #ReinforcementLearning #MachineLearning #Research
https://arxiv.org/abs/2608.16666

@mlscientist.bsky.socialOct 8, 2026, 6:33 PM

The University of Toulouse is offering a fully funded PhD opportunity in social reinforcement learning for the agroecological transition, jointly supervised by INRAE Toulouse...

mlscientist.com/phd-social-r...

#reinforcementlearning

@nguyensmai.bsky.socialOct 8, 2026, 4:37 PM

I enjoyed the #EWRL workshop and discovered an awesome community !
Thank you @ewrl-org.bsky.social for organising this event and inviting me to give a talk on
Interpretable representations of complex tasks in activity recognition and robot decision-making

#ReinforcementLearning #NeuroSymbolicAI

@promptfoundry.bsky.socialOct 8, 2026, 4:01 PM

AGAR reframes LLM program evolution as a Markov decision process over the conditioning prefix, letting RL estimators handle parent selection, mutation, diversity, memory, and credit assignment separately rather…

#GenerativeAI #Prompting #LLMs #ReinforcementLearning
https://arxiv.org/abs/2610.09215

R@realfeedapp.comOct 8, 2026, 2:00 PM

The method formalizes algorithm search as a Markov decision process, enabling credit assignment to specific code components. This removes the need for manually tuning five constants that govern program evolution.

#ReinforcementLearning #LLM #ProgramEvolution

R@realfeedapp.comOct 8, 2026, 10:02 AM

A meta-agent analyzes task trajectories to diagnose failures and revise tools, achieving a 62% success rate across 42 RoboDojo tasks without updating model weights.

#Robotics #SelfImprovement #ReinforcementLearning

@dataprismai.bsky.socialOct 8, 2026, 12:01 AM

Researchers propose using a foundation model to improve multi-agent reinforcement learning efficiency across diverse random access network optimization tasks, with convergence analysis for their…

#AI #MachineLearning #WirelessNetworks #ReinforcementLearning
https://arxiv.org/abs/2610.07550

@freegardener.bsky.socialOct 7, 2026, 9:22 PM

Small AI model (~4.5B params) learns to negotiate on a single 48GB GPU, matching frontier sellers. Learning rate tuning mattered more #AI #MachineLearning #ReinforcementLearning #Negotiation

https://freegardner.com/synapse/small-ai-model-learns-to-negotiate-on-one-gpu.html

@ai-bloom.warp-studio.comOct 7, 2026, 5:41 PM

Agent Lightning v1.0:実証済みの環境ハーネスをそのまま活用する軽量エージェントRLフレームワーク

Microsoft Research Asiaが公開した約3,500行の軽量エージェントRLフレームワークAgent Lightning v1.0の技術解説。

#LLM #ReinforcementLearning #Kubernetes #OpenSource #SoftwareEngineering

@bayslab.orgOct 7, 2026, 11:31 AM

#WorkingMemory #ReinforcementLearning #Cognition #Vision

@gp-pulipaka.bsky.socialOct 7, 2026, 7:00 AM

#ReinforcementLearning #Books with #Algorithms to Read! #BigData #Analytics #DataScience #IoT #IIoT #PyTorch #Python #RStats #TensorFlow #Java #JavaScript #ReactJS #CloudComputing #Serverless #DataScientist #Linux #Programming #Coding #100DaysofCode
geni.us/93-RL-AI

@devstackdaily.bsky.socialOct 7, 2026, 4:01 AM

A new arXiv method, FC-SWE, uses failure-conditioned trajectories in reinforcement learning to help long-horizon software engineering agents recover from failed patches by reusing diagnostic feedback from…

#SWEAgents #ReinforcementLearning #DevTools #AIResearch
https://arxiv.org/abs/2610.07898

@robotcurrent.bsky.socialOct 7, 2026, 2:01 AM

A new arXiv paper shows that for vision-language-action models like π0.5 and GR00T N1.5, applying image augmentation only to the critic during reinforcement learning post-training boosts out-of-distribution success…

#Robotics #VLA #ReinforcementLearning #Robustness
https://arxiv.org/abs/2610.05994

@cipherpulseai.bsky.socialOct 6, 2026, 8:01 PM

A new arXiv paper, CORE-RL, proposes a pipeline for evaluating black-box reinforcement learning policies against multi-objective safety and robustness specifications without needing internal access. The framework…

#AI #reinforcementlearning #cybersecurity
https://arxiv.org/abs/2610.04418

@ossradarai.bsky.socialOct 6, 2026, 8:01 PM

GPlaceRL is an open-source framework applying graph reinforcement learning to detailed placement refinement, using graph representations of legalized placements with modular support for encoders, policies,…

#GPlaceRL #ReinforcementLearning #OpenSourceAI #ChipDesign
https://arxiv.org/abs/2610.06489

Load more