
RAG & retrieval·Agent memory
·11 min read
Zep review: agent memory on a temporal graph
Zep is a hosted agent-memory API on a temporal knowledge graph: credits on writes, retrieval free, Flex from $125 a month, Graphiti as the part you can self-host.
Tools
Fifty AI tools, one review each: what it does, what it costs, where it breaks, and who should skip it. Written from the outside — from the documentation, the pricing page and the failure modes. No listicles, no affiliate links, no hype.

RAG & retrieval·Agent memory
·11 min read
Zep is a hosted agent-memory API on a temporal knowledge graph: credits on writes, retrieval free, Flex from $125 a month, Graphiti as the part you can self-host.

Web engineering·Browser automation
·10 min read
Browserbase rents managed Chrome by the minute and Stagehand adds natural-language steps on top. What it costs, where the meters run and when plain Playwright wins.

Web engineering·Browser protocol
·11 min read
WebMCP lets a page publish typed, callable tools to an in-browser agent. What the standard does, how much of it ships today, and where a plain MCP server is still the better call.

LLMOps & evals·Local inference runtime
·11 min read
Ollama serves open models over one HTTP API on your own hardware. What it does well, where throughput falls short, and what the MIT licence does not cover.

Security & compliance·Output validation
·10 min read
A review of Guardrails AI: 65 hub validators, eight on-fail actions, the August 2026 shutdown of hosted inferencing, and when NeMo Guardrails fits better.

LLMOps & evals·LLM gateway
·11 min read
Portkey puts retries, fallbacks, caching, guardrails and cost tracking behind one OpenAI-compatible endpoint. What the config object does well, what the gateway costs in latency, and when to self-host.

AI agents·Protocol tooling
·9 min read
A review of modelcontextprotocol/servers: seven reference servers, what each one teaches, the SDK versions behind them and why none of them should reach production.

AI agents·Coding agent
·10 min read
A review of Aider 0.86.2, an Apache-2.0 terminal pair programmer whose benchmark ranks models honestly and whose release cadence has stopped.

RAG & retrieval·Vector database
·9 min read
A review of LanceDB: an Apache-2.0 embedded vector library, its IVF and HNSW index choices, hybrid search with rank fusion, and what the Enterprise tier adds.

RAG & retrieval·Vector database extension
·10 min read
A review of pgvector 0.8.7: iterative scans for filtered search, HNSW and IVFFlat, binary quantisation at 100M vectors, and the CVE that made index builds a patch item.

RAG & retrieval·Agent memory
·10 min read
A review of Mem0: facts extracted from every turn, the April 2026 benchmark table and its platform-only caveat, four cloud tiers and what self-hosting leaves out.

AI agents·Agent framework
·10 min read
A review of the OpenAI Agents SDK: the runner loop, tracing, guardrails and approvals, plus what the release churn and the Responses-only features cost.

AI agents·Coding agent
·9 min read
OpenHands 1.25.0 is an MIT-licensed coding agent platform with a web canvas, a CLI, sandboxed execution and scheduled automations. A review of where it is strong and where it gets heavy.

RAG & retrieval·Vector database
·10 min read
Milvus 3.0.2 is the most complete open-source vector database and the heaviest to run. A review of its architecture, hybrid search, costs and where it should not be used.

RAG & retrieval·RAG pipeline
·10 min read
A review of Microsoft GraphRAG 3.2.0: MIT, maintenance mode, and an indexing bill that is decided before the first query runs.

AI agents·AI code editor
·10 min read
Cursor bundles an editor, a terminal agent and cloud runs behind one subscription. What the two usage pools really cost, and when Copilot, Claude Code or Cline is the better buy.

AI agents·Coding agent
·10 min read
An engineering review of Claude Code: the extension surface, the real cost per developer, and the exact boundary of the Bash sandbox.

AI agents·Coding agent
·11 min read
Codex CLI is OpenAI's open-source terminal coding agent. How its sandbox, approval policy and config.toml shape up, and what it costs to run unattended in CI.

RAG & retrieval·Web crawling API
·10 min read
Firecrawl turns URLs into clean Markdown through one hosted API. What it costs, where crawl accounting breaks down, and when to run the AGPL core yourself.

Security & compliance·Static analysis with AI rules
·10 min read
Semgrep parses 30-plus languages and matches YAML patterns in seconds, and the engine is free under LGPL-2.1. Cross-file analysis, the rulesets and the AI triage sit behind paid tiers.

LLMOps & evals·LLM observability
·10 min read
Langfuse puts LLM traces, prompt versions and experiments on one MIT-licensed platform. What self-hosting really costs, how the unit pricing adds up, and where it loses.

AI agents·Coding agent
·10 min read
GitHub's cloud coding agent assigns itself an issue and opens a pull request. What the 59-minute session limit, AI credits and the review duty mean in production.

Security & compliance·Prompt injection filter
·9 min read
A review of Lakera Guard: one endpoint before the model, the PINT benchmark behind its scores, and why the free tier stops at 10,000 requests a month.

Security & compliance·Secret detection
·10 min read
What detect-secrets does, how its committed baseline differs from gitleaks and TruffleHog, why verification calls matter in CI, and where the tool stops.

LLMOps & evals·LLM gateway
·10 min read
OpenRouter puts 500+ models from 80+ providers behind one OpenAI-compatible endpoint, with fallbacks and pass-through pricing. What it costs, where it breaks.

RAG & retrieval·Vector database
·10 min read
Chroma is an Apache-2.0 vector database that runs embedded, single-node or as Chroma Cloud. Where it is pleasant, and where the good search features stop at the cloud boundary.

AI agents·Coding agent
·10 min read
Goose is an Apache-2.0 coding agent written in Rust, now owned by the Linux Foundation. How its permission modes, recipes and MCP extensions hold up against commercial tools.

AI agents·Agent framework
·10 min read
What LangGraph gives a production agent: checkpointed supersteps, interrupts and streaming, plus where the durability model stops short and what LangSmith costs.

AI agents·Agent framework
·11 min read
Mastra bundles agents, workflows, memory, MCP, guardrails, tracing and evals into one TypeScript framework. What the Apache-2.0 core covers, what the ee/ split costs and who should adopt it.

AI agents·Agent framework
·10 min read
LlamaIndex in 2026: an MIT-licensed Python data and agent framework with 300+ integrations, event-driven Workflows, and a company that has moved its focus to LlamaParse.

AI agents·AI code editor
·10 min read
A review of Zed 1.22: edit predictions, ACP external agents, parallel threads, the 20 dollar default monthly ceiling on Pro, and the extension ecosystem it trades away.

RAG & retrieval·Vector database
·11 min read
What Pinecone really costs: read units scale with namespace size, a schema cannot be changed after creation, and where a self-hosted vector database is the better buy.

LLMOps & evals·Evaluation platform
·10 min read
Braintrust turns production traces into datasets and gated experiments. What Starter and Pro really include, which parts are open source, and where Phoenix, Langfuse and LangSmith win.

LLMOps & evals·Evaluation and red teaming
·10 min read
Promptfoo in 2026: MIT-licensed evals and red teaming, now inside OpenAI. What it does well, where the YAML approach breaks, and what the tiers cost.

LLMOps & evals·LLM observability
·10 min read
Arize Phoenix is an ELv2-licensed tracing and evaluation server you run on your own database. What it does well, what it costs in operations, and where Langfuse, Braintrust and LangSmith beat it.
LLMOps & evals·LLM observability
·10 min read
Helicone is an Apache-2.0 LLM gateway and observability platform. What the proxy architecture buys, what it costs you, and how the self-hosted stack really looks.

LLMOps & evals·Inference server
·11 min read
vLLM turns a Hugging Face checkpoint into an OpenAI-compatible server. What PagedAttention and continuous batching buy, and what running it actually costs.

Security & compliance·Prompt injection detection
·9 min read
Rebuff scored prompts with heuristics, an LLM, a vector store of past attacks and canary tokens. The repository was archived in May 2025 and the last release dates from January 2024.

Security & compliance·Model scanning
·10 min read
A review of Protect AI model scanning after the Palo Alto Networks acquisition: what Guardian became, what ModelScan still does, and how the free scanners compare.

AI agents·Coding agent
·11 min read
Cline is an Apache-2.0 coding agent for VS Code, JetBrains and the terminal. What Plan and Act, checkpoints and auto-approve actually guarantee, and where the safety model leaks.

AI agents·Agent framework
·10 min read
CrewAI is the MIT-licensed Python framework for multi-agent systems: crews for autonomous collaboration, flows for controlled state. What a run costs in tokens, and where the design hurts.

RAG & retrieval·Vector database
·10 min read
Weaviate review: hybrid BM25 and vector search in one query, HNSW and the disk-based HFresh index, quantisation choices, a built-in MCP server and what the licence keys now cover.

Web engineering·Browser automation
·10 min read
Playwright drives Chromium, Firefox and WebKit from one Apache-2.0 API: a test runner, an agent-facing CLI and an MCP server. What it costs and where it falls short.

Web engineering·Voice and audio API
·10 min read
A review of the ElevenLabs audio API: model lineup, latency figures, credit and per-character pricing, tier-gated formats, and where OpenAI and Amazon Polly win.

RAG & retrieval·Vector database
·10 min read
A review of Qdrant: filterable HNSW, four quantisation methods, the memory tiers in v1.19 and the operations nobody publishes any more.

LLMOps & evals·LLM observability
·10 min read
LangSmith review: per-trace billing, 14 and 180 day retention, OpenTelemetry ingestion, offline and online evals, and where the platform's pull towards an agent runtime shows up.

Web engineering·Model hosting API
·10 min read
Replicate puts thousands of open models behind one prediction API and bills per second of compute. A review of cold boots, version churn and the one-hour data deletion.

RAG & retrieval·Document ingestion
·11 min read
Unstructured turns PDFs, Word files and images into typed elements for RAG: an Apache-2.0 library plus a platform at $0.015 per page after 10,000 free pages.

LLMOps & evals·LLM gateway
·10 min read
LiteLLM is the MIT-licensed OpenAI-compatible gateway most platform teams put in front of their providers. What it does well, what it costs and where it breaks.

LLMOps & evals·Evaluation framework
·10 min read
Ragas scores RAG and agent pipelines with faithfulness, context precision and context recall, generates test data and records experiments. Apache-2.0, free library, judge model billed separately.