So sad to hear that the iconic, trailblazing, Apollo 11 moon landing software engineering lead, Margaret Hamilton, is no longer with us. Love this image of her, standing next to a print out of the software that landed Buzz Aldrin and Neil Amstrong on the moon! 👩🔬🔭🛰️🧪
ℹ️: www.bbc.co.uk/news/article...
The one and only Michael Jordan was interviewed last week by the French newspaper Libération. Here is an English translation. Enjoy!
www.di.ens.fr/~fbach/MJord...
Some biggies in this release: ■ The CatEncoder can be super useful ■ Caching makes exploratory work so much more productive And many places got faster (always nice to have) or easier to use. Enjoy more powerful and easier learning with dataframes :)
🧠 TextEncoder is now LLMEncoder -- clearer name, same transformer embeddings
📊 DataOp reports now link to your source code and show docstrings -- can be generated without executing calculations
⚡ The SessionEncoder is now up to 15x faster
🔍 describe_transformations() -- can now print a plain text explanation of what TableVectorizer and Cleaner did to each column
💾 Persistent caching for DataOps -- pipeline steps can now be cached across runs
🆕 CatEncoder -- OneHotEncoder + TargetEncoder, two encoders combined in a single column transformer to deal with columns with many infrequent classes
🫂13 new contributors helped with this release!
Also in this release: faster construction of deep DataOps, ToCategorical for numeric columns, and more!
✨ Skrub 0.11 has been released ✨
Full changelog:
skrub-data.org/stable/CHANG...
Heads-up: skrub 0.11 requires Python ≥ 3.11 and scikit-learn ≥ 1.5.2.
Highlights in thread ⤵️
Tomorrow, I'm on the Vanishing Gradients podcast to talk about agentic science with @hugobowne.bsky.social, Shipra Arora from Bain & Company, and Luca Fiaschi from @pymc-labs.bsky.social
Online live panel, with QA!
🗓️ Fri Oct 2, 1pm CEST / 7am EST
➡️ Register: numfocus-org.zoom.us/webinar/regi...
Super excited about our Survey with Lihu Chen on the Role of Small Models in the LLM Era
This topic is crucial in today's world
direct.mit.edu/coli/article...
More details about those changes and other fixes in the changelog:
The README now also includes a section about how to empirically assess the subtle semantic variations of various BLAS and OpenMP runtimes:
The documentation in the README of the project has been updated to explain how to achieve this.
Mitigating this problem is needed to unlock the full value of free-threading Python, especially for @scikit-learn.org workloads that often nest BLAS calls (via NumPy, SciPy or PyTorch) and OpenMP calls (via Cython) under Python level threads (typically via joblib).
Oversubscription problems typically happen when nesting BLAS or OpenMP calls under Python threads: naively spawning 10 Python threads that themselves spawn 10 BLAS threads each results in 100 starving threads on a 10 cores CPU.
This release includes several contributions by itamarst.hachyderm.io.ap.brid.gy from @quansight.com. in collaboration with myself & others at @probabl.ai. It provides tools to inspect the semantics of native threadpools in various environments so as to be able to mitigate oversubscription problems.
