How we extended Apache DataFusion to execute one query across many machines
Learn how Datadog built Distributed DataFusion to scale interactive Apache DataFusion queries across multiple machines.
Company engineering blog
Check out the Datadog Engineering blog to discover posts written by Datadog’s very own engineers.
6 posts in the last 90 days · newest 1 Oct
Learn how Datadog built Distributed DataFusion to scale interactive Apache DataFusion queries across multiple machines.
Learn how we made Python profiling async-aware while reducing profiler overhead and preserving context across asyncio tasks.
Learn how Datadog improved Rust tracing by building an opinionated OpenTelemetry-based library to help ensure consistent sampling, propagation, and trace quality at scale.
Learn how Datadog built gitretriever to handle 20× more CI Git traffic while maintaining low latency and reducing backend CPU usage.
Learn how the Datadog APM team improved Java startup performance by encoding a prefix trie as a JVM string constant.
Modern Java profilers often rely on unsupported JVM internals for accurate CPU profiling. Here’s how engineers from Datadog, SAP, Amazon, and the OpenJDK community helped bring a new CPU profiling event to JDK 25.
Learn how Datadog helps ensure data completeness at scale, enabling accurate alerts and safer automated decisions across distributed pipelines.
Using AI-assisted refactoring, we migrated our live routing brain to a relational model, safely validating changes against live production traffic.
During a reliability gameday, Datadog engineers discovered that their PostgreSQL clusters couldn’t safely fail over. Here’s how the team redesigned them for high availability using Patroni and synchronous replication.
By combining stacked LLM evaluations with tool-driven investigation, we scaled malicious code detection from pull requests to dependency packages without sacrificing accuracy or cost control.
Learn how we embed widget metadata into screenshots using invisible, resilient watermarks, enabling self-describing visualizations at scale.
Find out how we built a scalable evaluation platform for Datadog’s Bits AI SRE agent that replays real incidents, detects regressions, and measures agent performance across production scenarios.
When a high-volume upsert doubled disk writes, Datadog engineers traced the issue to Postgres WAL behavior and rewrote the query to eliminate hidden costs.
Learn how Datadog detected and resolved issues from hackerbot-claw, an AI-powered automated attack campaign.
We share lessons learned building Datadog’s MCP server, from designing agent-friendly tools and managing context windows to using queries instead of raw data retrieval.
Get an inside look at shrinking large Go binaries in the Datadog Agent through dependency and linker analysis.
Learn field-tested lessons for eBPF-powered workload protection.
Learn how Datadog scaled eBPF-powered file monitoring to handle more than 10 billion kernel events per minute while preserving full detection coverage.
Discover how Datadog engineered a scalable Change Data Capture (CDC) platform to replicate data across systems in near real time—reducing search latency by 87%, increasing availability, and powering diverse, multi-tenant use cases across the company.
Learn how Datadog’s SDLC Security team built an LLM-powered system to detect malicious pull requests at scale—without sacrificing developer velocity.