Glossary
-

Unlocking Apache Spark Performance: Three Open-Source Tools We Use at CERN
Unlocking Apache Spark Performance: Three Open-Source Tools We Use at CERN Apache Spark is incredibly powerful, but anyone who has worked with it long enough knows the feeling: Why is this job suddenly slower today? Why are executors running out of memory? Why is one stage taking 90% of the runtime? What exactly is Spark…
-

CERN PGDay 2026 is here!
CERN PGDay 2026 is here! After a successful first edition in 2025, CERN PGDay returns in 2026 as a regular gathering for PostgreSQL users and enthusiasts in Suisse Romande (western Switzerland). Co-organized by CERN and SwissPUG, the event offers a chance to connect, share ideas, and exchange experiences in the vibrant Geneva region — home…
-

Why I’m Loving Spark 4’s Python Data Source (with Direct Arrow Batches)
Why I’m Loving Spark 4’s Python Data Source (with Direct Arrow Batches) TL;DR: Apache Spark 4 lets you build first-class data sources in pure Python. If your reader yields Arrow RecordBatch objects, Spark ingests them with reduced Python↔JVM serialization overhead. I used this to ship a ROOT data format reader for PySpark. A PySpark reader…