Zero or null. Mathematically correct or practically useful. Marco Gorelli digs into a real design trade-off between Polars and SQL, and shares an opinion on which one gets it right.
Full talk: buff.ly/L0duaRh

Pandasの`read_csv`でメモリ爆発して絶望したこと、一度はあるよね? Polarsに変えるだけで処理速度が跳ね上がり、メモリも激減します。 import polars as pl df = pl.scan_csv("data.csv").filter(pl.col("a") > 0).collect() Lazy評価を使えば巨大ファイルも落ちずに爆速処理。もうPandasのメモリエラーに怯える必要はありません! #Python #Polars
Zero or null. Mathematically correct or practically useful. Marco Gorelli digs into a real design trade-off between Polars and SQL, and shares an opinion on which one gets it right.
Full talk: buff.ly/L0duaRh
📝 "Pills dataset - Part 1"
Exploring Rust for data science with the Pillbox dataset.
👤 Jennifer HY Lin (@jhylin.bsky.social)
🔗 https://jhylin.github.io/Data_in_life_blog/posts/09_Pills/Rust_polars_pills_ws.html
#pyladies #python #oldiebutgoodie #dataanalyticsprojects #pillsdatasetseries #polars
Polars ETL script: reads CSV, applies your transform, writes Parquet. 210 lines, intermediate. Beats Pandas on memory-heavy pipelines. https://www.valtersit.com/python/polars-dataframe-etl-pipeline/ #python #data #polars
Small differences in engine semantics create big shifts in query output.
Marco Gorelli breaks down the operational differences between Polars and traditional SQL.
Read Marco's technical breakdown & migration tips: labs.quansight.org/blog/polars-...
#Polars #SQL #DataScience #OpenSource #Narwhals
Polars 2.0 RCが示した「派手さより、まず壊れにくさ」
https://papoo.work/doc/19f9368897bc197d
#polars #データ分析 #python #ストリーミング処理 #api変更
Everyone rushes to compare Polars and SQL on speed and syntax. Marco Gorelli points out that almost nobody is comparing the semantics, the actual mental models behind each one.
Full talk from PyData London: buff.ly/uInE29g
#Polars #SQL #PythonData #DataEngineering #Quansight
Polars 2.0 RCが示した「派手さより、まず壊れにくさ」
https://papoo.work/doc/19f9368897bc197d
#polars #データ分析 #python #ストリーミング処理 #api変更
Polars 2.0 defaults to the streaming engine now instead of in-memory - about 5x faster on most queries. It also stops guaranteeing row order on joins and group_by unless you pass maintain_order=True yourself.
The kind of default that passes every test and reorders your data weeks later.
高速データ処理ライブラリ「Polars」の次期メジャー版、2.0の先行リリースが登場。クエリエンジンの一新により、さらなるパフォーマンス向上とAPIの最適化が期待されます。データエンジニア必須のアップデートとなりそうです。 #Polars #DataEngineering