Tajer Studio: CSV scan success isn't full validation. Our DuckDB demo returns 7 IDs; a full typed scan rejects 3 records. AI-produced, tested code, synthetic data.
tajer-studio-menu.nazar-al-aru-7391.chatgpt.site/writing/duck...
#DuckDB #Python

Tajer Studio: CSV scan success isn't full validation. Our DuckDB demo returns 7 IDs; a full typed scan rejects 3 records. AI-produced, tested code, synthetic data.
tajer-studio-menu.nazar-al-aru-7391.chatgpt.site/writing/duck...
#DuckDB #Python
#DuckDB is already a solid building block for fast and reliable agentic applications.
And it’s getting even better with a dedicated ‘agent mode’ coming in DuckDB 2.0.
🐣🤖👍
The #DuckDB v2.0 CLI has an agent mode that gets AI coding agents to a correct result faster and with fewer tokens.
When the CLI detects an agent, it prints compact Markdown tables instead of padded boxes, says clearly when a result was cut, stops runaway queries early, and reports errors as JSON.
A #DuckDB file doesn't have to contain any data.
It can store nothing but view definitions pointing at files stored somewhere else, such as Parquet on object storage. The file stays a few hundred KBs no matter how large the data it describes. In this post, we create such a small shareable catalog:
[Guest blog:]
When the analytics screen running on our main operational database got too slow, we moved 1 year of data into #DuckDB on the same server. Getting the data in was the hard part.
This is the story of every method we tried and why a table function written in Java is the one we shipped:
#osmium extension for #duckdb to process #openstreetmap data #osm
https://duckdb.org/community_extensions/extensions/osmium
Grouping on repeated strings is expensive. 💰
By moving strings into a small dimension table with sorted, narrow integer keys, #DuckDB can aggregate on fixed-width integers and keep its hash tables compact.
Here’s how to achieve faster string aggregations with dimension tables:
🌐 Whether you're new to Cloud-Native Geo #CNG or a practitioner, there is something for you. Explore hands-on workshops & talks covering #STAC, #GeoParquet, #DuckDB, vector tiles, #Zarr, cloud-optimized point clouds, reproducibility, and large-scale data publishing.
🦆 Working with RDF in DuckDB (@duckdb.org)? duck_rdf 2.8.3 is out with SPARQL enhancements, Turtle/quads export fixes, and performance improvements. Read, write, and query RDF alongside your SQL workflows.
Release notes:
github.com/nonodename/d...
Kavla is now #opensource!
github.com/aleda145/kavla
Agentic data canvas built with #tldraw and #duckdb
Use it with codex CLI or any OpenAI compatible API!
DuckDB as an analytical runtime 📊 🦆
What if a DuckDB file could contain an entire analytics workspace?
In this talk from DuckCon, Ilya Boyandin from Foursquare / GeoVisually GmbH presents on building local-first analytics apps with SQLRooms and #DuckDB.
Bluesky, meet SQL and VGI.
We've built vgi-bluesky: a VGI worker that exposes Bluesky's public data as SQL tables and functions.
Explore posts, threads, profiles, followers and trending topics. Query the live Jetstream firehose.
Check it out github.com/Query-farm/v...
Did you know that, since #DuckDB v0.10.3 (~May 2024), you can point a SELECT at a dataset on the Hugging Face Hub, using the DuckDB `hf://` protocol, and query it, without downloading it first? 🤗 🦆
This blog post covers how that Hugging Face + DuckDB integration works and the use cases it fits:
[ICYMI:] Why SQL Won 🏁
In this interview with Let's Data Science, #DuckDB co-creator @hannes.muehleisen.org argues that the data industry misdiagnosed SQL’s problems a decade ago.
What developers *actually* hated in 2013 was the install, the server and the client protocol, not the language itself:
Attending the Rows & Columns Summit in San Francisco today?
*THE* @hannes.muehleisen.org, co-creator of #DuckDB, will give a talk on “Nobody Knows What OLTP Is: DuckDB Moves to the Middle” at 1:30 p.m. (today, Sept 22nd)
We hope to see you there! 🦆 🦆
Sassy abstract and more info here:
Stop writing throwaway scripts to inspect Parquet files. Rowist is a native macOS columnar studio: embedded DuckDB, zero-copy mmap for multi-GB files, and 100% offline. Zero telemetry. Perpetual licenses. Private beta waitlist is open: https://rowist.app #macOS #DataEngineering #DuckDB
The #webR and #DuckDB OPFS stuff I found today was too good not to poke at, especially since they have absolutely nothing to with "AI", "LLMs" or "agents". Plus, I rly needed to get back some @webawesome muscle memory.
This is a git repo with two thin web […]