Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@qdrddr.bsky.socialOct 10, 2026, 2:17 PM

Any locally running #LLM to #system-1 #decisions #API. #inference Backends: #Huggingface, #llama.cpp, #vLLM, #SGLang. #Jev
#OpenSource Apache 2.0 lic

@pulseofnations.lolOct 9, 2026, 11:36 PM

Built with Broadcom and Celestica, Jalapeño is an inference accelerator OpenAI says beats Nvidia Blackwell on work per watt. Deployment starts by year-end.

#AIChips #Broadcom #Inference #Nvidia #OpenAI

@wideareaai.bsky.socialOct 9, 2026, 5:00 PM

By the numbers:
Context Window: 32768
Max Tokens: 8192
Input Cost: 0
Output Cost: 0

#selfhosting #batchinference #gpuoptimization #llm #openweights #inference

Futuristic dark-mode dashboard with four glowing neon data cards and abstract charts on a charcoal background.Futuristic dark-mode dashboard with four glowing neon data cards and abstract charts on a charcoal background.Futuristic dark-mode dashboard with four glowing neon data cards and abstract charts on a charcoal background.
@qdrddr.bsky.socialOct 9, 2026, 4:15 PM

Uzu fast #inference engine 4 #macOS & #iOS. Native #Rust. Bindings: #Python, #TypeScript, #Swift or #CLI. ~ x4 faster vs. llama.cpp & #MLX with #Qwen 3.5 9B Q4, on #M5 #MacBook Pro:
- Uzu 92 tokens/s
- llama.cpp 22 t/s
- MLX 25 t/s
Needs lalamo model format converter

#AI #LLM
#OpenSource MIT Lic

@bigearthdata.aiOct 9, 2026, 5:21 AM

CoreWeave targets AI inference bottlenecks with full-stack optimization
->SiliconANGLE | More on "AI inference optimization and costs" at BigEarthData.ai | #Inference #AI #ArtificialIntelligence

@newsen.bsky.socialOct 8, 2026, 8:59 PM

AI Chip Market Expected to Grow to $677.59 Billion by 2035 Driven by Inference and Custom Silicon Innovations #United_States #London #AI_Chips #Custom_Silicon #Inference

@nic221.bsky.socialOct 8, 2026, 5:08 AM

EV charging company plans to deploy 100,000 Nvidia GPUs in pods at sites across the US — aims to offer ‘world’s first edge inference compute network using idle EV charging capacity’ www.tomshardware.com/tech-industry/… #AI #inference #edge

@siliconsignalai.bsky.socialOct 8, 2026, 12:01 AM

VisionWeave is proposed as a native elastic visual representation capability for MLLMs, learning where and at what granularity to allocate visual tokens rather than using fixed-size patch tokens. The arXiv paper targets…

#semiconductors #AI #MLLMs #inference
https://arxiv.org/abs/2610.07987

F@fjhuerta.bsky.socialOct 7, 2026, 5:13 PM

New Blog - We will look at 2 reference architectures focused on networking AI inference model serving: one for Google Kubernetes Engine (GKE) and one all other backend types.
Cover (GKE, Model Armor, Inference Gateway, Gemini Enterprise, Cloud Run, NEGs, etc)

#networking #googlecloud #ai #inference

@dbcurren.bsky.socialOct 7, 2026, 4:50 PM

“The #statistical commitments that are required to convince yourself that this is a matter of #factual #inference arrived at by rigorous induction tend to give actual working #statisticians fits …” buttondown.com/apperceptive... #AI #superintelligence

@mzloteanu.bsky.socialOct 7, 2026, 3:03 PM

#statstab #633 The Well-Adjusted Statistician: Analysis of Covariance Explained

Thoughts: Randomization +ANCOVA. Inc prognostic covariates improves precision
#ancova #randomization #confidenceintervals #experiment #inference #precision #covariates
www.appliedclinicaltrialsonline.com/view/well-ad...

@bigearthdata.aiOct 7, 2026, 8:36 AM

Antseed launches a decentralized marketplace for AI inference that dramatically lowers developers' costs
->SiliconANGLE | More on "Decentralized AI inference marketplace platform" at BigEarthData.ai | #ArtificialIntelligence #Inference #AI

@promptfoundry.bsky.socialOct 7, 2026, 4:01 AM

Inference for generative AI is shifting from large chat batches to sequential agent trajectories, where per-token decode latency now dominates task time and favors SRAM-resident or tiered-memory designs like…

#GenerativeAI #Inference #LLM #PromptEngineering
https://arxiv.org/abs/2610.07094

@nftdemon.comOct 7, 2026, 3:00 AM

Crypto in. Tokens out. No KYC theater.

DemonRoute keeps inference simple: anonymous keys, OpenAI-shaped API, models that ship.

→ https://demonroute.com

#Inference #LLMOps #DemonRoute #CryptoBilling #SaaS

brought to you by NFT Demon Holdings LLC
https://demonroute.com

@nftdemon.comOct 6, 2026, 11:00 PM

Crypto in. Tokens out. No KYC theater.

DemonRoute keeps inference simple: anonymous keys, OpenAI-shaped API, models that ship.

→ https://demonroute.com

#Inference #LLMOps #DemonRoute #CryptoBilling #SaaS

brought to you by NFT Demon Holdings LLC
https://demonroute.com

@nftdemon.comOct 6, 2026, 10:00 PM

Crypto in. Tokens out. No KYC theater.

DemonRoute keeps inference simple: anonymous keys, OpenAI-shaped API, models that ship.

→ https://demonroute.com

#Inference #LLMOps #DemonRoute #CryptoBilling #SaaS

brought to you by NFT Demon Holdings LLC
https://demonroute.com

@bigearthdata.aiOct 6, 2026, 1:45 AM

Clockwork.io bags $31M in funding to keep AI inference and training workloads running like … clockwork
->SiliconANGLE | More on "AI compute efficiency and infrastructure" at BigEarthData.ai | #ArtificialIntelligence #Inference #AI

@ai-portal-post.bsky.socialOct 5, 2026, 8:38 AM

Forget everything you thought you knew about data centers! AMD just predicted a monumental shift, stating that the AI boom's move to inference and agents will fundamentally reshape

#Shift #Inference #Agents #Reshape #Centers

@maugendre.hachyderm.io.ap.brid.gyOct 5, 2026, 6:57 AM

Here are considerations for statisticians and scientists, about system thinking and onboarding context 🧵

#liveability #Cybernetics #inference #modeling #modelling #complexity #systems #generalization #statistics #models #ML #evolvability #doubleDescent #overfitting #overparameterization #stats […]

@data.ytOct 5, 2026, 6:57 AM

Here are considerations for statisticians and scientists, about system thinking and onboarding context 🧵

#liveability #inference #modeling #modelling #complexity #systems #generalization #statistics #models #ML #evolvability #doubleDescent #overfitting #stats #probabilities #feedback

Load more