InfoQ Live Logo

The Software Architects' Newsletter
September 2026
View in browser

Welcome to the InfoQ Software Architects' Newsletter! Each month, we bring you essential news and insights on emerging patterns and technologies from industry peers.

This month, we focus on "Building Data and AI Platforms". Technologies, patterns, and practices from this topic span the entire "diffusion of innovation" graph in last year's "InfoQ Software Architecture and Design Trends Report".

We examine how architects design data systems that balance governance with accessibility, including data mesh, decentralized ownership models, and platform engineering approaches that make data infrastructure self-service. As organizations move beyond monolithic data warehouses toward composable, product-oriented data platforms, understanding the architectural trade-offs in scalability, discoverability, and cross-team enablement is becoming a core concern.

News

From Retrieval to Reasoning: Building Production-Ready Agentic AI Systems with Knowledge Graphs

Cassie Shum explains why knowledge graphs are a critical foundation for agentic systems. Moving beyond basic RAG, she explains four practical architectural patterns: context bundling, decision provenance, code as truth, and agent visibility. Cassie demonstrates an engineering harness built on a knowledge graph to streamline feedback loops, optimize token usage, and maintain system reliability.

In a complementary podcast, Michael Stiefel spoke with Scott Hanselman about developing new software engineers when artificial intelligence agents do most of the work junior developers are trained on.

Beyond Offset Lag: Computing Time in Queue for Apache Hudi Data Lake Pipelines at Petabyte Scale

Srikanth Mamidala explains how Twilio improved monitoring of its Kafka and Apache Hudi pipelines by complementing traditional offset lag with a “time-in-queue” metric that measures how stale unprocessed data actually is. The approach provides clearer visibility into data freshness and enables more meaningful pipeline SLAs and early warnings.

Your Next DSL Author Is a Language Model

Irakli Betchvaia argues that LLMs struggle with domain-specific languages because their syntax is poorly represented in training data. He proposes "Typed Domain Grounding", embedding DSLs in familiar languages such as Kotlin or TypeScript, where type systems and compiler feedback can catch hallucinated syntax and help models iteratively repair their output.

Platform Engineering in the Age of AI

In this recording of a recent InfoQ Live webinar, the panelists explain how platform teams adapt to support AI-assisted engineering, highlighting which capabilities belong in the platform. They discuss trade-offs between standardization and developer autonomy, while sharing strategies to manage AI tooling, security guardrails, and shifting workflows.

Context Engineering at LinkedIn: How We Built an Organizational Context Layer for AI Agents with MCP

Ajay Prakash discusses how LinkedIn overcomes AI agent limitations in large codebases. He explains Contextual Agent Playbooks and Tools, built on Model Context Protocol (MCP), which serves procedural memory, code search, and runbooks directly to coding agents. Prakash shares architectural details and operational guardrails that deliver a twenty percent productivity boost with zero loss in reliability.

From S3 to GPU in One Copy: Rethinking Data Loading for ML Training

In this QCon London presentation recording, Onur Satici explains how Vortex, an open-source columnar file format under the Linux Foundation, revolutionizes high-throughput data loading. He details how cascading lightweight encodings, layout-based segment pruning, and zero-copy memory pipelines eliminate CPU/NVMe bottlenecks, streaming S3 data straight to GPUs at speeds up to sixty gigabits per second without upfront data reprocessing.

Case Study

Implementing Durable Workflows on Postgres Without an External Orchestrator

Raman Varma explores how PostgreSQL can provide durable workflow execution without introducing a dedicated orchestrator such as Temporal or AWS Step Functions. Drawing on the architecture of Kestrel Workflows, Varma shows how familiar database primitives can handle the core requirements of workflow orchestration: SELECT ... FOR UPDATE SKIP LOCKED provides concurrent work claiming, primary key constraints enforce idempotent step checkpoints, and leases with a periodic sweeper enable recovery when workers fail.

The approach also supports long-running sleeps and human approval steps by persisting waits as database state, allowing workflows to survive application restarts or Kubernetes rescheduling. Because executions and checkpoints remain relational data, operational visibility becomes straightforward SQL rather than requiring the export of workflow data into a separate analytics system.

The architecture is not universally applicable. Varma highlights connection pressure, table churn, and dispatch latency as considerations at larger scale, while suggesting that a dedicated engine such as Temporal is preferable for workloads requiring thousands of workers or extremely low dispatch latency. For I/O-heavy automation with modest concurrency, however, the article argues that reusing an existing highly available PostgreSQL deployment can reduce infrastructure, security, and operational complexity.

This content is a short summary of a recent InfoQ article by Raman Varma, "Implementing Durable Workflows on Postgres Without an External Orchestrator".

To get notifications when InfoQ publishes content on these topics, follow "AI, ML & Data Engineering", "Big Data", and "Database" on InfoQ.

Architecture Decisions Across Platforms, AI, and Delivery

Where should a platform standardize, and where should teams retain flexibility? What should an AI agent be allowed to access? What catches a coding agent's mistakes before its changes reach CI?

InfoQ's live online certification programs give you five weeks to work through decisions like these using QCon frameworks, facilitator context, and perspectives from experienced engineers at other companies.

Upcoming programs include:

Explore the October cohorts.

Missed a newsletter? You can find all of the previous issues on InfoQ.

The IIoT PostgreSQL Performance Envelope - Sponsored by Tiger Data

Industrial IoT systems can move from a fast pilot to a data-scale problem surprisingly quickly. As sensor counts and historical data grow, ingest, storage, and query workloads can expose architectural limits that weren't visible at proof-of-concept scale. This guide explores how to define a PostgreSQL performance envelope for IIoT workloads, anticipate capacity and storage constraints, and understand the tradeoffs of different scaling approaches. Learn how to evaluate database architecture for production-scale IIoT before performance limits become an operational problem.

Download the guide “The IIoT PostgreSQL Performance Envelope,” sponsored by Tiger Data

Upcoming Events

About InfoQ

Senior software developers rely on the InfoQ community to keep ahead of the adoption curve. One of the main reasons software architects and engineers tell us they keep coming back to InfoQ is because they trust the information provided and selected by their peers.

We've been helping software development teams adopt new technologies and practices for over 20 years through InfoQ articles, news items, podcasts, tech talks, trends reports, and QCon software development conferences.

We hope you find this newsletter useful. If not, you can unsubscribe using the link below.

Unsubscribe

Forwarded email? Subscribe and get your own copy.

Subscribe

 

Follow InfoQ on:

You have received this email because you subscribed to "The Architects' Newsletter". To stop receiving the Architects' Newsletter, please click the following link: Unsubscribe

- - -

C4Media Inc. (InfoQ.com), 705-2267 Lake Shore Blvd. West,
Toronto, Ontario, Canada, M8V 3X2