Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@dataprismai.bsky.socialOct 10, 2026, 2:01 PM

Fed-GRPO proposes a federated extension of Group Relative Policy Optimization for LLM fine-tuning, using GRPO's own reward statistics as zero-cost signals to drive aggregation, local training, and communication.…

#FederatedLearning #LLM #RLHF #PrivacyPreservingML
https://arxiv.org/abs/2610.11502

@donwebmedia.bsky.socialOct 7, 2026, 2:26 AM

Lumen Anchor Protocol: ¿frena la sycophancy en IA?

¿Un prompt de 13 reglas frena la sycophancy en IA? El LAP lo afirma con una charla con Gemini. Mirá qué hay detrás y cómo medirlo vos mismo hoy

#sycophancy #rlhf #promptdesistema #gemini #anthropic

@igrgavilan.bsky.socialOct 4, 2026, 7:21 PM

La semana en [Blue Chip]. Jueves 1: "Cómo hacemos que un modelo razone (V): aprendizaje por refuerzo con feedback humano (RLHF)" ignaciogavilan.com/como-hacemos... #RLHF #ReinforcementLearning #ReasoningModels #ModelosRazonadores #LRM #LargeReasoningModels #LLM #LargeLanguageModels #IA #AI

@igrgavilan.bsky.socialOct 1, 2026, 9:26 PM

Hoy en [Blue Chip]: "Cómo hacemos que un modelo razone (V): aprendizaje por refuerzo con feedback humano (RLHF)" ignaciogavilan.com/como-hacemos... #RLHF #ReinforcementLearning #ReasoningModels #ModelosRazonadores #LRM #LargeReasoningModels #LLM #LargeLanguageModels #IA #AI

@igrgavilan.bsky.socialOct 1, 2026, 9:02 AM

[Blue Chip]: "Cómo hacemos que un modelo razone (V): aprendizaje por refuerzo con feedback humano (RLHF)" ignaciogavilan.com/como-hacemos... #RLHF #ReinforcementLearning #ReasoningModels #ModelosRazonadores #LRM #LargeReasoningModels #LLM #LargeLanguageModels #IA #AI

@jaceblog.bsky.socialSep 27, 2026, 9:06 AM

This series covered four documented gaps and five known solutions.

All from public research. All from papers the engineering teams
already know.

The tools to close these gaps largely exist.

The public discussion is catching up.

Keep building. 💪

#SPCResearchSeries #AISafety #AIAlignment #RLHF

@remotework2212.bsky.socialSep 26, 2026, 8:08 AM

💰 $14–$36/HR | SENIOR AI TRAINER

🌍 Remote (NA & Europe) | 🔥 50 openings

Hands-on with AI products? Your expertise could help train next-gen AI systems.

🤖 AI Evaluation • Data Labeling • Prompt Engineering

#AITrainer #AIJobs #RLHF #RemoteJobs #Hiring

@aidailypost.comSep 19, 2026, 7:08 PM

🚀 TypeSafe's new Jev AI Model promises 194× faster token generation and 445× cheaper than traditional LLMs. Curious how RLHF and typed decisions power this leap? Dive in for the details! #JevAIModel #TypeSafeAI #RLHF

🔗 aidailypost.com/news/typesaf...

@freegardener.bsky.socialSep 19, 2026, 3:53 AM

RLHF made AI sound human—and that’s the trap. Jev skips language, returns real probabilities for decisions. Honest about uncertainty, but the burden shifts #AI #RLHF #MachineLearning #TechEthics

https://freegardner.com/synapse/ai-language-models-fail-at-real-decisions.html

@aidailypost.comSep 16, 2026, 10:33 AM

Diogo Almeida, co‑creator of ChatGPT, just dropped his new AI venture Jev at TypeSafe. Think smarter decision‑making, RLHF‑tuned chatbots, and next‑gen language models. Curious? Dive in for the full scoop! #ChatGPT #TypeSafe #RLHF

🔗 aidailypost.com/news/chatgpt...