Skip to content
HomeInsights

Insights

Ideas for building
better software.

Engineering perspectives on architecture, product development, AI, and the everyday decisions behind useful software.

From the engineering desk

Browse practical articles or follow the latest technology updates.

Subscribe via RSS →

Curated links from external sources — not 360Softy original articles.

ExternalCloud
Cloudflare Changelog

Workers AI - Planned model deprecations on Workers AI

We are refreshing the Workers AI model catalog to make room for newer releases. Please update your apps to remove references to the models listed below before the deprecation date. Recommended replacements @cf/zai-org/glm-4.7-flash — fast multilingual model with multi-turn tool calling and coding capabilities. @cf/google/gemma-4-26b-a4b-it — efficient open model with vision and tool calling. @cf/moonshotai/kimi-k2.6 — capable tool-calling and vision model for agentic workloads and coding. For pr

Workers AI
Cloudflare ChangelogRead original
ExternalFrontend Development
Vercel Blog

Chat SDK adds web adapter support

You can now build chat UIs that connect to Chat SDK with the new . Build in-product assistants, support agents, or any other browser-based chat experience.web adapter Define the bot on your server: Then stream replies to the browser with a preconfigured hook:@ai-sdk/reactuseChat Read the to get started, browse the , or build your own .documentationdirectoryadapter Read more

Vercel BlogRead original
ExternalFrontend Development
Vercel Blog

Chat SDK now supports conversation history

Chat SDK now supports cross-platform conversation history through the new and options. User transcripts persist across every , allowing the same user to keep their message history wherever they message your bot.transcriptsidentityplatform adapter exposes four methods, backed by one of the :bot.transcriptsofficial state adapters Read the to get started, or try one of the .documentationtemplates Read more : persist an inbound message or a bot replyappend : return entries chronologically wi

Vercel BlogRead original
ExternalDatabase
Redis Blog

What’s new in two: April 2026 edition

Welcome to “What’s new in two,” your quick hit of Redis releases you might have missed in the past month. If you blinked, you missed it—so here’s the recap. We’re covering the latest developments from April and expanding on what I covered in our lates...

Tech
Redis BlogRead original
ExternalAI
NVIDIA Technical Blog

Achieving Peak System and Workload Efficiency on NVIDIA GB200 NVL72 with Slurm Block Scheduling

NVIDIA GB200 NVL72 introduces a fundamentally new way to build GPU clusters by extending NVIDIA NVLink coherence across an entire rack. This design enables... NVIDIA GB200 NVL72 introduces a fundamentally new way to build GPU clusters by extending NVIDIA NVLink coherence across an entire rack. This design enables exascale performance, but it also changes the assumptions that many scheduling systems were built on. As a result, “rack-scale locality” becomes a hard constraint. When workloads cross

NVIDIA Technical BlogRead original
ExternalAI
NVIDIA Technical Blog

Model Quantization: Post-Training Quantization Using NVIDIA Model Optimizer

Model quantization is an effective method to reduce VRAM usage and improve inference performance on consumer devices such as NVIDIA GeForce RTX GPUs. By... Model quantization is an effective method to reduce VRAM usage and improve inference performance on consumer devices such as NVIDIA GeForce RTX GPUs. By lowering computational and memory requirements while preserving model quality, quantization helps AI models run more efficiently in resource-constrained environments. This post walks through

NVIDIA Technical BlogRead original

Let’s start with a conversation

Tell us what you’re working on.

An idea, a challenge, or a system that needs to work better. We’ll help you understand the next step.