Glossary

  • Unlocking Apache Spark Performance: Three Open-Source Tools We Use at CERN

    Unlocking Apache Spark Performance: Three Open-Source Tools We Use at CERN

    Unlocking Apache Spark Performance: Three Open-Source Tools We Use at CERN Apache Spark is incredibly powerful, but anyone who has worked with it long enough knows the feeling: Why is this job suddenly slower today? Why are executors running out of memory? Why is one stage taking 90% of the runtime? What exactly is Spark…

  • CERN PGDay 2026 is here!

    CERN PGDay 2026 is here!

    CERN PGDay 2026 is here! After a successful first edition in 2025, CERN PGDay returns in 2026 as a regular gathering for PostgreSQL users and enthusiasts in Suisse Romande (western Switzerland). Co-organized by CERN and SwissPUG, the event offers a chance to connect, share ideas, and exchange experiences in the vibrant Geneva region — home…

  • Why I’m Loving Spark 4’s Python Data Source (with Direct Arrow Batches)

    Why I’m Loving Spark 4’s Python Data Source (with Direct Arrow Batches)

    Why I’m Loving Spark 4’s Python Data Source (with Direct Arrow Batches) TL;DR: Apache Spark 4 lets you build first-class data sources in pure Python. If your reader yields Arrow RecordBatch objects, Spark ingests them with reduced Python↔JVM serialization overhead. I used this to ship a ROOT data format reader for PySpark. A PySpark reader…