Sign in to view Rahul’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Rahul’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Bellingham, Washington, United States
Sign in to view Rahul’s full profile
Rahul can introduce you to 10+ people at Anthropic
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
119K followers
500+ connections
Sign in to view Rahul’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Rahul
Rahul can introduce you to 10+ people at Anthropic
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Rahul
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Rahul’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Activity
119K followers
-
Rahul Patil shared thisClaude Fable 5.1 and Claude Mythos 5.1 are available today. They're the most advanced models we've built for coding and knowledge work, and their research capabilities give an early look at how AI will contribute to science. Mythos 5.1 designed protein binders with a nearly 50% hit rate across 12 targets, confirmed in the lab, where 10 to 15% is typical today. On three targets, its binding affinities were ten times higher than the best designs from Adaptyv Bio's protein design competitions. It also sped up seven open-source biology models by up to 2.5x in a few days, cutting GPU costs on genome-wide analyses by 30 to 60%. That work usually takes a team of performance engineers weeks, and we plan to open-source the optimizations soon. Fable 5.1 scored 52.6% on Terminal-Bench-Science (Fable 5 hit 24.7%) and 55.8% on Terminal-Bench 4.0. Millennium used it to find the cause of a one-in-a-million crash their engineers had chased for years. Our automated audit also finds it better aligned than its predecessor across most metrics, with less reward hacking and less willingness to ignore explicit constraints. The system card has the full picture, including where we still fall short. Fable 5.1 also improves on some of the feedback we've heard from you, with more improvements coming in future models. Cost. Cache reads drop 75% to $0.25 per million tokens for customers on the API. That's about 25% cheaper for typical workloads and up to 45% for highly agentic work. Standard pricing stays at $10/$50. Safeguards. Fable 5.1 can now find vulnerabilities in source code, and our cyber safeguards intervene about 60% less often. Biology safeguards fire 85% less often on benign requests, and we've started a trusted access program with the US government for life scientists who need the full model. Data retention. Enterprise Frontier Safeguards, rolling out this fall, keep your data in cloud infrastructure you control, with the privacy of zero data retention. Eligible customers get ZDR on 5.1 starting now. Writing. It's clearer and more concise, and early testers noticed. Canva called the writing "the standout" in the new model. Fable 5.1 matches or beats Fable 5 at Low or Medium effort at much lower cost, so revisit your effort settings as you test it out, not just your model string. Read the blog for more: https://lnkd.in/gnGznuaW
-
Rahul Patil shared thisThrilled to welcome Ankur Sinha to Anthropic to lead enterprise and verticals engineering!! Ankur will be focused on how Claude delivers for major enterprises, regulated industries, and the public sector. His experience leading product and engineering at enterprise scale will help us keep up with the demand for Claude at work. He joins us from Remitly, where he was Chief Product and Technology Officer.Rahul Patil shared thisI've joined the amazing mission and team at Anthropic! Moving from transforming lives with trusted financial services that transcend borders (Remitly's mission) to ensuring the world safely makes the transition through transformative AI (Anthropic's mission), a few things stay constant. The impact on human lives is what makes the work worth doing, and trust and safety are critical to doing it well. I'm super stoked to be part of this and to help carry it forward. I'll be leading our Enterprise & Verticals areas. On the Enterprise side, we're focused on helping our enterprise customers transform how they work with AI. On Verticals, we're pushing to unlock real value in new domains - Life Sciences and Healthcare, Financial Services and Legal, Cyber and Public Sector - and continuing to push the envelope on what's possible with our models and products. I spent my first week at HQ and came away even more energized. There's been a lot of curiosity about how Anthropic works and where we see AI going. Among many other resources including our newly launched Anthropic Academy, this (https://lnkd.in/gsYscErQ) is a good representation of the progression as we understand it. I learned an enormous amount over the past few years serving as the Chief Product and Technology Officer at Remitly in building trust with our customers and enabling our teams to progress towards our mission, and I'm eternally grateful to Remitly's customers and teams for that. Starting at Anthropic, I'm equally grateful to our customers for trusting us as we collectively navigate this transition with AI, and to our teams working tirelessly to advance the mission. If you want to work on the hardest problems in enterprise AI - reliability at scale, safety in regulated domains, and enabling real value and scaled deployments in Life Sciences and Healthcare, Financial Services, Legal, Cyber and Public Sector - we're hiring across Enterprise & Verticals. Come build with us - https://lnkd.in/gi2qfH6q (Apply directly here if you see roles that resonate)
-
Rahul Patil shared thisLaunching Claude Opus 5! It comes close to Claude Fable 5 across many domains at half the price. It's our fourth Claude 5 model in under two months, and the one I expect to be a solid daily driver for most people. Opus 5 is the new state of the art on knowledge work and on Frontier-Bench, where it more than doubles Opus 4.8's score at a lower cost per task. On CursorBench 3.2 at max effort, it lands within 0.5% of Fable 5's peak score at half the cost per task. And it beats every other model on performance at a given cost at high, xhigh, and max effort. It's much smarter at autonomous work than any benchmark suggests. It verifies its own work and iterates carefully until it succeeds, fixing root causes instead of symptoms, and building its own checks when the right ones don't exist. In one benchmark task, Opus 5 had to rebuild a machine part as a 3D CAD model from a drawing it had no way to view. So it wrote its own computer vision pipeline to pull the geometry from the raw pixels, and solved the task repeatedly, while no competing model solved it once in five attempts. Cristian Rivera, an engineer at Stripe, put it well: "Over one weekend, I gave it a chief-of-staff role over my dev environments: it built its own monitor, drove each box, and pulled me in only for the judgment calls." Opus 5 is for the work your teams run all day: coding, agents, knowledge work. Now with near-frontier capability at $5/$25 per million tokens, unchanged from Opus 4.8. Fable 5 remains the model for your most ambitious work: the days-long autonomous projects nothing else can take on. If you're on any Opus today, this is a great upgrade. Same API, same price, swap the model string. More in the blog: https://lnkd.in/gjX3jN3P
-
Rahul Patil shared thisWe just published details on Fable's cyber safeguards, plus a proposed Cyber Jailbreak Severity framework. This is a shared way to score jailbreaks by capability gain and discoverability. We welcome feedback and critique at cyber-safeguards@anthropic.com, and we've launched a HackerOne program where security researchers can submit potential cyber jailbreaks they discover in Fable 5 for our review. HackerOne: https://lnkd.in/gjeiSe_n Link to detailed blog: https://lnkd.in/ge2fT2Vx
-
Rahul Patil shared thisClaude Sonnet 5 is available today. Over the past year, the clearest capability gains have come from our largest models. This is the biggest jump we’ve ever seen for a Sonnet model: roughly Opus 4.7 class on coding and agent tasks. For builders deciding what to run in production, it's the best cost to capability ratio in the Claude lineup. Our goal is for you to have the best model at every size. Opus for the hardest problems, Sonnet for the work your teams run all day, Haiku when speed and volume matter most. Sonnet 5 is the biggest step we've taken on the middle of that curve. It's also a drop-in upgrade. Same API, same price, same speed targets as Sonnet 4.6. Swap the model string and you get better results immediately on multi-file code changes, long agent runs, and browser automation, with a 1M token context window and 128K output. Sonnet 5 is live today in Claude Code, and on the Claude API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry. $3 per million input tokens and $15 per million output, with introductory pricing of $2/$10 through August 31. Read more: https://lnkd.in/g4zv7wBw
-
Rahul Patil shared thisTag, Claude's it! Today we’re launching Claude Tag (@Claude) and it’s changing the way engineering teams ship, again. Less synchronous, more async. Less work done alone in a private editor, more done in the open alongside teammates. And increasingly, work that closes its own loop. An agent picks up a problem, breaks it down, and comes back with something finished. @Claude is more than a tool, it works alongside you in Slack. It keeps up with the channels you point it at and decides when to chime in, schedules its own follow-up work, and reacts to events from tools like GitHub. It's multiplayer, so anyone in a channel can prompt it and pick up where someone else left off. System administrators specify which tools and information Claude should have access to, in which channels, so it stays scoped to the channels, data, and repos you choose, with every action auditable. And it runs on its own credentials (never your team's). This is how Anthropic builds now. 65% of our product team’s new code is created by our internal version of @Claude, and our engineers are shipping on the order of 8x the code per quarter they did over the last few years. We've said before that we're handing more of our own development to Claude. This is beginnings of what that looks like in practice. Launching in beta today in Slack for Claude Enterprise and Team customers, with more surfaces soon. If your team already builds with Claude, tagging @Claude in is an important next step. More in the blog: https://lnkd.in/g-9ZcjVz
-
Rahul Patil shared thisClaude Fable 5 is available today! It's a new moment for AI: a Mythos-class model, the most capable class of systems we've built, now safe for general use. It's already changed how we work internally, and I'm excited to see what you all do with it. Every request runs past safety classifiers trained to detect misuse in cybersecurity and biology. When one triggers, your request is answered by Opus 4.8 instead. More than 95% of sessions never see a fallback, and 1,000+ hours of external red-teaming produced no universal jailbreak. In terms of benchmarks, Fable 5 reached 80.3 on SWE-bench Pro (Opus 4.8 scores 69.2), 88 on Terminal-Bench 2.1. State-of-the-art on nearly every coding benchmark we tested. But the benchmarks undersell how truly capable it is. Fable holds quality deep into long, hard problems where most models degrade. It verifies its own work. It catches what others miss, things like root-cause bugs that no other model had surfaced. Base44 found it "much deeper and better at one-shotting full apps"; at Genspark it came out #1, winning head-to-head against every model they tested. Internally, writing code stopped being the slow part a while ago — Anthropic engineers on average shipped 8x as much code per quarter as they did compared to 2021-2025 — Fable pushes the bottleneck further toward verification and review. We're excited to make all of that available today for every use case outside bio and cyber. For API customers, here's how we've imagined fallbacks: pass a fallbacks parameter and the Messages API retries any blocked turn on Opus 4.8 server-side — even mid-stream, keeping the partial output. We think of this as a graceful handoff between models, and we'll be iterating on the design with the community. Moments like this are worth doing right. We're making sure it's safe, but the classifiers may be annoying at times. They're tuned conservatively, and false positives will keep coming down. Read more here: https://lnkd.in/giBEAAcP
-
Rahul Patil shared thisWe just shipped Claude Opus 4.8 and dynamic workflows! Claude Opus 4.8 is the most capable model we've put out and the best you can build on right now, outside the Mythos-class systems we're still testing under Project Glasswing. SWE-bench Pro went from 64.3 to 69.2. But the improvement I keep coming back to is honesty. Opus 4.8 is about 4x less likely than 4.7 to let a flaw in its own code slide past unremarked. It tells you what it's unsure of instead of dressing up thin progress as finished work. For anyone whose agents run with real oversight cost, that's worth more than another point on a leaderboard. That compounds on big, long-running jobs, which is why I'm most excited about dynamic workflows, launching in research preview today. Claude plans the work, fans out across hundreds of parallel subagents in a single session, and verifies its own output before handing it back, with 4.8 letting those agents run longer before they report. This is aimed at the work that used to take a quarter and a working group: codebase-scale migrations, sprawling refactors, and bug fixes across hundreds of thousands of lines, graded against the test suite you already trust. Available in Claude Code for Enterprise, Team, and Max. One practical note for fellow builders: Opus 4.8 defaults to "high" effort, and that's the right setting for most work. There's an "xhigh" level for the genuinely hard, long-running tasks. It's strong, but it's token hungry, so reach for it deliberately. We've made the headroom for it on both sides: higher rate limits for subscribers, and fast mode now 2.5x the speed at 3x lower cost ($10/$50) for API customers. Standard pricing is unchanged at $5/$25. Available today. More in the blog: https://lnkd.in/ggSctZyc
-
Rahul Patil shared thisTalk to any CTO standing up AI in production and the same operational questions come up: who's the data processor, how does this plug into our identity stack, and can we keep procurement simple. Today we're launching the Claude Platform on AWS in general availability, a first of its kind offering for Anthropic. AWS customers get the full Claude API through the AWS account they already have. Authentication runs through AWS IAM. Audit through CloudTrail. Billing comes through a single AWS invoice and retires against committed spend. Every new Claude Platform feature and beta ships on AWS the same day it ships on our native API, including Claude Managed Agents, the advisor strategy, code execution, skills, and the MCP connector. For teams with regional data residency requirements, Claude on Amazon Bedrock remains available with AWS as the data processor. https://lnkd.in/ge_Cjj-kIntroducing the Claude Platform on AWS | Claude by AnthropicIntroducing the Claude Platform on AWS | Claude by Anthropic
-
Rahul Patil liked thisRahul Patil liked thisThis week is my last at Stripe. I joined in 2018 to work on Developer Productivity, focused on helping Stripe’s engineers have the most productive time of their careers. That started with the inner loop, trying to make the developer experience as tightly integrated as possible. I truly believe that Stripe has some of the best developer tools in the industry. I’m incredibly proud of what we built, but much more grateful for the people I got to build it with. Stripe is genuinely full of people who care deeply about their work, hold themselves and each other to a high bar and are just great people to spend your days working alongside. I’ve learnt an enormous amount from them, especially my managers Aaron Spinks, Rahul Patil and Will Larson. Thank you to all the current and former Stripes who made the last eight years so memorable. I feel incredibly fortunate to have spent them here.
-
Rahul Patil liked thisHighly recommend checking this blog out for how best to use Claude Tag for Data Analytics. https://lnkd.in/gV7RU_gR https://lnkd.in/gb29T-85Self-service data analytics in Slack: how Anthropic deploys Claude Tag for ad-hoc questions | Claude by AnthropicSelf-service data analytics in Slack: how Anthropic deploys Claude Tag for ad-hoc questions | Claude by Anthropic
-
Rahul Patil liked thisRahul Patil liked thisWe've rebuilt the Claude in Chrome side panel as a Claude Cowork session, with the same conversation history, skills, and connectors as Claude on desktop, web, and mobile. Sessions live with your account, not on any single device. Start a task in a browser tab and pick it up on your desktop or phone, with the whole conversation saved to your history. The new side panel is on the Max and Team plans today, and rolling out to the Pro plan in the coming weeks. Enterprise plan admins can enable it for their orgs. Browser agents can be tricked by instructions hidden in a page. We build defenses against this, and we still recommend a few habits of your own: https://lnkd.in/eSyFtHh3 Give it a try: claude.com/chromeThe Claude in Chrome side panel is now a Claude Cowork sessionThe Claude in Chrome side panel is now a Claude Cowork session
-
Rahul Patil liked thisRahul Patil liked thisWe’re building a custom silicon team at Anthropic. If these roles sound like you, reach out! Silicon Engineer https://lnkd.in/drcqbBGh Hardware Systems Architect https://lnkd.in/dbMiq7nk
-
Rahul Patil liked thisRahul Patil liked thisOver the last year, we’ve interviewed thousands of people about AI, including a study of 81,000 Claude users on what they want from the technology (which was one of my favorite projects we’ve done at Anthropic). Today, we’re building on that by sharing some of their voices — and the hardest questions they asked us. We won’t get the benefits of AI without addressing the hard questions. You can share your own here: https://lnkd.in/eGxCU_7P
-
Rahul Patil liked thisRahul Patil liked thisBig day! Claude Cowork is coming to web and mobile, so Claude can keep working while your computer is closed. This is a major update to Cowork. It combines the power of giving Claude access to your context, an advanced loop for long-running tasks, and the convenience of not needing your laptop to be open. Given the size of the change, it's rolling out over the next couple of weeks, starting with our Max users. For more details, check out the blog post: https://lnkd.in/gZdXxZxU
Experience & Education
-
Anthropic
***
-
*****
***** ******
-
******
***
-
********** ** **********
****** ** ******** ************** ***** undefined undefined
-
-
******* ***** **********
*** ******** *******
-
View Rahul’s full experience
See their title, tenure and more.
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
or
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View Rahul’s full profile
-
See who you know in common
-
Get introduced
-
Contact Rahul directly
Other similar profiles
-
Chris Tava
Chris Tava
I leverage data, machine learning, artificial intelligence, cloud computing, and agile methodologies to build software that integrates AI/ML into workflows to reduce administrative burden and deliver value.<br><br>With over 25 years of experience in software engineering, I have a proven track record of creating and scaling software for B2C and B2B businesses across various domains. <br><br>That being said, I love solving problems for healthcare! It's very satisfying to help clinicians and patients live their best lives.<br><br>I have a passion for applying machine learning and AI to business applications that make an industry-wide impact. <br><br>I hold certifications in Machine Learning from Coursera and Data Science and Machine Learning: Making Data-Driven Decisions from MIT.<br><br>Am currently studying for a Masters in Computer Engineering at Dartmouth College. This program has several machine learning courses: Machine Learning, NLP, Computer Vision and Deep Learning. Its also covers the hardware side with: Embedded Systems, Field Programmable Gate Arrays. Here is a link to the program: https://engineering.dartmouth.edu/graduate/meng/online-computer-engineering
8K followersNew York City Metropolitan Area
Explore more posts
-
Jookwang Jung
Fasient • 106 followers
OpenAI's CFO Sarah Friar published a piece laying out how the company thinks about delivering more intelligence at lower cost, and the framing is worth sitting with for a minute. The argument is that gains compound across the entire stack: chip efficiency, compute infrastructure, model improvements, and product delivery all reinforce each other. It's not one lever pulling harder, it's multiple layers improving in parallel. That compounding logic is actually what makes the economics of AI hard to predict from the outside. If you only track model benchmarks or headline API pricing, you might miss that a quieter efficiency gain two layers down is what actually changes the unit economics for builders and enterprises. Having a CFO make this case publicly is also a bit of a signal in itself. The conversation is shifting from "look what the model can do" toward "here's why the cost curve keeps moving" which is the kind of narrative that matters when you're trying to close large enterprise deals or justify continued infrastructure investment. Whether the compounding holds at the scale OpenAI is targeting is still an open question. But the framework they're describing, full-stack efficiency rather than model-only improvements, is a reasonable way to think about where durable cost reductions actually come from. https://lnkd.in/gUQBTAan
1
-
Sahand Sojoodi
I write about building useful… • 5K followers
My highlights from PyAI Conf - Mar 10 in SF 𝐀𝐈-𝐠𝐞𝐧𝐞𝐫𝐚𝐭𝐞𝐝 𝐨𝐩𝐞𝐧 𝐬𝐨𝐮𝐫𝐜𝐞 “𝐬𝐥𝐨𝐩” 𝐢𝐬 𝐧𝐨𝐰 𝐚 𝐫𝐞𝐚𝐥 𝐦𝐚𝐢𝐧𝐭𝐚𝐢𝐧𝐞𝐫 𝐭𝐚𝐱 This came up repeatedly: low-signal AI PRs/issues/reviews are increasing triage load for maintainers. Some ideas discussed were creating a "credit score-like system" for contributor reputation management. This has been on everyone's mind, and even outside of PyAI, Pete Steinberger (OpenClaw) has been very vocal about this and Daniel Stenberg (curl) has documented similar pressure at maintainer level. • Pete Steinberger post: https://lnkd.in/dWYgy9dU • Daniel Stenberg, Death by a thousand slops: https://lnkd.in/d9zhu4Su • Daniel Stenberg, The end of the curl bug bounty: https://lnkd.in/d2NtryCQ 𝐌𝐨𝐧𝐭𝐲 (Samuel Colvin / Pydantic) looks like a useful middle-ground runtime The Monty demo and discussion was strong because it���s intentionally minimal + fast, not trying to be a full VM/container replacement. It’s a Python interpreter in Rust designed for safe execution of model-generated code with explicit boundaries to host functions. That gives more flexibility than strict tool-calling, while staying more controlled than open-ended sandbox compute. • Monty repo: https://lnkd.in/dm4tzTJH • Pydantic deep dive: https://lnkd.in/dcqdSQ3P 𝐥𝐚𝐭.𝐦𝐝 - 𝐀 𝐤𝐧𝐨𝐰𝐥𝐞𝐝𝐠𝐞 𝐠𝐫𝐚𝐩𝐡 𝐟𝐨𝐫 𝐲𝐨𝐮𝐫 𝐜𝐨𝐝𝐞𝐛𝐚𝐬𝐞, 𝐰𝐫𝐢𝐭𝐭𝐞𝐧 𝐢𝐧 𝐦𝐚𝐫𝐤𝐝𝐨𝐰𝐧 Yuri’s https ://lat.md stood out as actionable: it creates a markdown-based knowledge graph for codebases with cross-linking, reference checks, and agent-friendly prompt expansion/search. This is interesting for onboarding + architecture navigation + reducing context loss in larger repos. • Repo: https://lnkd.in/du2chFrf 𝐁𝐚𝐮𝐩𝐥𝐚𝐧 - "Git for Data" Ciro’s Bauplan section was compelling because it maps software-dev primitives (branches/commits/transaction-like execution) to data workflows. Their transactional pipeline model is especially important: run in isolated branch, merge on success, keep main untouched on failure. That aligns well with agentic workflows where reproducibility, rollback, and auditability matter. • Bauplan: https://lnkd.in/dhjrwxMQ • Git for Data: https://lnkd.in/dPtDNM4n • LLM quick start: https://lnkd.in/d-wV6d7x What's New In 𝐅𝐚𝐬𝐭𝐀𝐏𝐈 Also worth adding: Sebastián's “What’s new in FastAPI for AI” highlighted that FastAPI is becoming increasingly AI-native in practice, improving dev UX via pyproject.toml-based setup and VS Code/Cursor extensions. He demoed new production-friendly streaming patterns (JSON Lines, bytes, and especially SSE) for real-time LLM responses; also the Pydantic return-type performance, making FastAPI feel optimized for modern AI backends: https://lnkd.in/dcXJSr4Z
65
9 Comments -
Kate Kruizenga
Phero Collective • 5K followers
Chat GPT / Claude / Gemini is not a reliable source of benchmark data. In the last 6 months I've heard: 🩺 The average spend on healthcare is only $X at Series A... 🌴 Did you know we don't have to offer PTO? 💰 Most ML engineers in SF make $Y. 📈 Everyone has plans with mega backdoor Roth functionality now! Reality? 🙈 Health Spend: Data source was *three* founders on a Reddit thread. We talked about how health plans are rated, how we use PEOs to offer stronger benefits at early stage, and how we need to make decisions about offering reasonable, cost-effective coverage to *our* group as part of a comprehensive total rewards strategy. Bonus: Brokers have high-fidelity data sets on your peer group! 🫠 PTO: Legally, yes. Practically this is noise that tells you to achieve goals by controlling how your employees work. If there are issues with velocity or goal attainment, we look at the dozens of other levers to inspire and reach outputs that are more effective than throttling or taking away PTO. 🤑 Salary: It's a smattering from Blind, Levels, Glassdoor... Self-reported salary data is not high-fidelity. $100K-500K salary ranges on job postings are not helpful. What's is? Professionally managed, scrubbed and validated (including levels, not titles) salary benchmark data. Pave is still free if you connect for T1 benchmarks, or you can pay $5-10K for a data set to do this right. It's an investment in fair comp, better offer accept rates, and retention. 🤦♀️ 401K: They sure don't outside very big tech, and we aren't Google. We pulled from someone's Hampton conversation cut-paste. Your 10-person startup doesn't need a retirement plan this complex; let's talk about your distinctives and holistic Total Rewards philosophy. GPT can be a thought partner to a busy leader or professional, but it's also highly likely to bring you incomplete data that reinforces the position you ask it to validate, and then help you workshop how to sell it. ❓ Ask non-leading questions, and ask for different points of view, risks, and pitfalls when sparring with LLMs. 💡 Seek out better, high-fidelity data. 🧠 Tap someone with technical expertise that can help you avoid unnecessary mistakes and get the outcomes you're actually aiming for. Fractional Chief Financial and People Officers offer 1-hour consults and can help you quickly cut through the noise and find a high-impact solution for *your* unique business. Thankful for the many fractionals, VC partners, and high-fidelity benchmark reports out there doing the work. Let's help founders build better foundations, scale intelligently, and unlock teams to do their best work! 🚀
36
12 Comments -
Vijoy Pandey
Cisco • 19K followers
We've been scaling in one direction ⬆️, with bigger models, more parameters, and more compute. That vertical progress will continue. But there's a second axis we've we've been ignoring, ➡️ outward. Today, one agent figures something out, that knowledge lives and dies with that agent. Every agent optimizes for its own goals. We need systems that optimize for emergent multi-agent goals. Connecting agents or APIs doesn't fix this - that's passing opaque data, not meaning. We need both vectors: scaling up AND scaling out. Sat down with Axios to unpack what that means and why it changes the path to distrbiuted superintelligence. Link in the comments.
54
1 Comment -
Ashu Garg
Foundation Capital • 44K followers
Arvind Jain sees context graphs emerging from enterprise search and observability. I agree on the destination and disagree on the starting point. Arvind - a fellow IITD alum who's devoted his career to enterprise search - has built a tremendous business at Glean, and his "capture the how, learn the why over time" framing is sharp. But the foundation needs to be different: 1 - Search sits in the read path, not the write path. By the time activity data reaches the index, the decision context - why an exception was granted, what inputs were weighed, who approved - has already been flattened. 2 - Inferring intent from observed patterns is fundamentally different from capturing decision context in the operational flow. One reconstructs; the other records. 3 - The ~80% accuracy Arvind cites on task understanding is impressive, but when agents affect customers, contracts, and compliance, you need an authoritative record, not a probabilistic inference. (The math of compounding errors is brutal - I’ve argued for 99%+ accuracy in past editions of my Substack.) -- We believe orchestration-layer startups - where the context graph is built by actually executing work - are better positioned: 1 - They sit at the point of decision. When an agent triages an escalation or approves a discount, it pulls from multiple systems and applies policies in real time. That's the moment to capture the decision trace, not reconstruct it later from activity signals. 2 - They create decision traces as first-class artifacts. The orchestration layer captures the full picture: what inputs were gathered, what policies applied, what exceptions were granted, and what state existed at the moment of decision. That's not inferred; it's captured - which means enterprises can answer "why did we do that?" definitively. 3- They can learn across customers. Arvind notes enterprise data can't be aggregated for privacy reasons. But orchestration-layer startups focus on bounded workflows - which means they can refine ontologies within a customer and across deployments without ever sharing raw data. -- PlayerZero, Maximor, and Oliv are great examples of this. PlayerZero builds the context graph by automating L2/L3 support - sitting at the intersection where code, config, infrastructure, and customer behavior collide. Maximor AI captures decision lineage by orchestrating finance workflows where reconciliation logic and exceptions actually live. Oliv AI builds it for sales - starting as a co-pilot that performs specific tasks (update CRM, send follow-up) and capturing the decision traces along the way. There will be multiple context graphs within each org, and enterprise search can certainly be a starting point for one of them. Glean's approach creates value for knowledge retrieval and understanding how work flows. The question is whether the authoritative record of decisions will emerge from observability or from orchestration. We're betting on orchestration.
221
48 Comments -
Kumar Chellapilla
Inception • 7K followers
Today we’re launching Mercury 2 (https://lnkd.in/gP8pZuDA), the fastest reasoning LLM and the first reasoning dLLM: ~1,000 tokens/sec, ~5× faster than leading speed‑optimized LLMs. What I’m most excited about isn’t just the speed benchmark line, but what this unlocks for building real AI systems: multi‑step agents that don’t stall, voice assistants that can reason inside tight latency budgets, and long-form coding loops that stay in flow. Diffusion changes the generation loop: Mercury works more like an editor iterating in parallel than a typewriter committing one token at a time. We’re hiring engineers across research, systems, infrastructure, and product—If you’re excited about working on frontier LLMs and care deeply about engineering quality, DM me—I’d love to connect (https://lnkd.in/g28qHbNV).
155
8 Comments -
Josh Clemm
Figma • 7K followers
Everyone's AI agents (or openclaws...) still need great context. We recently published an article sharing various lessons in Context Engineering while building Dash at Dropbox. We still believe building an index across all your 3rd party apps gets you much higher quality context (and far more reliable). You also don't have to just use the Dash product to access that index. Connecting the Dash MCP to your favorite AI app is one of the fastest ways to bring in a ton of your work context. You get connectors, content understanding, cross-app graphs, and search all in one. Other lessons I touch on: - Using knowledge graphs for cross-app intelligence - Challenges we faced with MCP tool calling - Promising wins using DSPy at scale - Heavy use of contextual LLMs as a Judge Let me know if folks have additional questions in the comments. https://lnkd.in/eqqnvN9F
122
8 Comments -
Rubén Domínguez Ibar
The VC Corner • 336K followers
When One Engineer Is 1000x At Sequoia Capital, early product risk comes down to people. In this clip, Pat Grady and Alfred Lin, with Jack Altman, explain why one exceptional engineer can define a company’s DNA. ▫️ ServiceNow went public with most of its core code written by Fred Luddy ▫️ Airbnb was largely built early on by Nate airbn alone These weren’t 10x engineers. They were closer to 1000x. That early technical ownership shapes everything that follows. I broke this down in the Sequoia Playbook, including 10 systems most firms never had the patience to build, here: https://lnkd.in/e9ubxeyM Where have you seen one builder change everything? Do you agree?
6
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content