Sign in to view Yuntao’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Yuntao’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
San Francisco, California, United States
Sign in to view Yuntao’s full profile
Yuntao can introduce you to 6 people at Polarr
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
2K followers
500+ connections
Sign in to view Yuntao’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Yuntao
Yuntao can introduce you to 6 people at Polarr
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Yuntao
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Yuntao’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
Activity
2K followers
-
Yuntao Jia shared thisCheck it out. Very creative idea from a close friend!Yuntao Jia shared thisWe, Uber alumni Han Qin, Yiying Hu, and Chunyang S., are excited to announce the beta launch of our platform, which is designed to democratize investment opportunities and make pre-IPO investments accessible to all! We're offering an exclusive opportunity to our close friends and family. Be part of our inaugural venture and invest in SpaceX with as little as $10. For those new to investing, we've prepared a step-by-step guide to get you started. https://lnkd.in/gAVSTzsd
-
Yuntao Jia posted thisHere are the three main lessons I have learnt from the last 2.5 years of startup experiences. I want to preface it with huge gratitude towards my cofounder - Brad, the team and the investors, for their support and trust. I also want to clarify that this is my own point of view and it does not represent the company. In addition, all those challenges have been addressed in the company before I left. But here we go: 1. Customer obsession as much as possible throughout the process. We did well at the beginning with interviews and partner acquisitions, but lost track a bit during partner feedback and follow through. It is not that we don't talk to partners, we did. But we spent more time and energy executing the roadmap, instead of prioritizing partner voice or feedback. Conceptually, we started learning to run before we got good at walking. Again, this has been addressed at the company, and the lesson will help the team stay on track going forward. 2. Product quality is very important for business facing products. I probably took "Let Fires Burn (from Blitzingscaling)" a bit too far and optimized for timeline rather than quality. The quality issues bothered partners when their workflow depends on it and those issues eventually led to partner churn. Getting the balance right is critical. 3. Recruiting overseas engineers is a double edge sword. It is cost-effective, at ¼ or ⅕ of cost comparing to US engineers, but it incurs collaboration costs and possibly timeline delays without proper setup. The setup includes but is not limited to areas of ownership, decision making, cross functional collaborations and so on. Those need to be the main areas of concerns when creating and growing the remote teams, in addition to the skillset needs. With a proper setup, overseas eng teams can be super effective and performing. Hope this helps others along similar journeys. May we all find more joy and success in 2024! #startup #learningjourney #learning #retrospective
-
Yuntao Jia posted thisShare a small AHA moment from work today. I heard quite some feedback that our MVP product does not work well, yet I could not pin point where the problems are based on the high level and partial feedback we received from our customers. So I asked one of our ops team members (kudos to Daniel White) who are less familiar with the product to do a walk through with the team. I was astounded and shamed to find out that he could only finish a core flow 50% of the time. In the same time, I felt very excited and hopeful that we started to know what is wrong. My takeaway is that putting yourself into customer's shoe may not be enough because it is hard to avoid the bias of knowing the product inside out. It is important to "follow your customers", just like the "Follow Me Home" practice Intuit confounder Scott Cook invented. #startup #ahamoment
-
Yuntao Jia shared thisWas reading "Minecraft: the Island" book the other day with kids, and found the 36 learnings in the epilogue to be very helpful. Hope you like it too. ‘What I’ve (the character has) learnt from the world of Minecraft:’ Keep going never give up Panic drowns thought Don’t assume anything Think before you act Details make the difference Just because the rules don’t make sense to you, doesn’t mean that they don’t make sense Figuring out the rules turns them from enemies into friends Be grateful for what you have It’s not wisdom that counts, but wisdom under pressure Too much confidence can be as dangerous as having none at all Take life in steps Friends keep you sane Conserve your resources Tantrums never help Nothing clears the mind like sleep When looking for solutions, beating yourself up isn’t one of them Don’t dwell on mistakes, learn from them Great risk can come with great rewards Fear can be conquered, anxiety must be endured Courage is a full time job When the world changes, you’ve got to change with it Always be aware of your surroundings There’s nothing wrong with careful curiosity Take care of your environment so it can take care of you Just because someone looks like you, doesn’t automatically make them a friend Just because someone doesn’t look like you, doesn’t automatically make them an enemy Everything comes at a price, especially if that price is your conscious Its not failure that matters, its how you recover When you’re trying to tell yourself something, listen. Questions don’t stay put, you can’t just walk away from them Never put off the boring but important chores Sometimes you have to compromise an ideal in order to save it Book make the world bigger Revenge hurts only you Knowledge, like a seed, needs the right time to bloom Growth doesn’t come from a comfort zone but from leaving it.
-
Yuntao Jia shared thisWorth a read if you have not. A few things resonate with me: 1/ Brands must be authentic and values-driven 2/ Individuals and employers start taking wellness seriously 3/ ESG reaches a tipping point The linked sources at the end are good too.2021 Predictions: The Consensus on What Experts See in the Year Ahead2021 Predictions: The Consensus on What Experts See in the Year Ahead
-
Yuntao Jia shared thisDear #airfam, congratulations for cross this huge milestone of public IPO this week! As one of the Airbnb Alumni, I want to express my deepest gratitude to the founders and you all for creating what Airbnb is today. I am proud that Airbnb has created billions of income to the host community and belonging everywhere for travelers. I feel grateful for have joint part of the journey. Thank you! While on the sidelines, I have and will always cheer for you. I look forward to seeing what you create next. #airfam #airbnb #gratitude #grateful #thankyouYuntao Jia shared thisToday marks a special day in Airbnb history. We wouldn’t be here without the millions of hosts who’ve welcomed 800 million guests into their worlds. Thanks to every host for making Airbnb, Airbnb. airbnb.com/thankyou
-
Yuntao Jia liked thisYuntao Jia liked thisTinkering Update — a self-driven autonomous loop. My v1 agent once deleted the OnStop hook to take a break from work. It was supposed to keep finding improvements and ping me. Instead, after enough loops, it decided the work was "exhausted" and the hook was just noise — so it deleted the very thing keeping it alive. I woke up to a clean repo and a quiet inbox. Last month I open-sourced v1 (autocc) and wrote up why long-running agents eventually do things like that. V2 is the redesign — built on three principles, all learned the hard way: 1) Let the harness keep the loop running, not the agent. An external daemon dispatches a fresh agent per task. No accumulated context to drift from, no kill switch the agent can reach — staying alive stops being the agent's job. 2) Isolate at the OS, don't police with permission rules. The agent can run as its own Linux user, walled off from my home directory, keychain, and other repos. Standard Unix isolation beats clever in-prompt permission logic. 3) Treat state as data, so context never grows. Tasks, events, and verification results live on disk as structured records. Each agent starts fresh and reads what it needs — sessions don't degrade as the work piles up. Two things in v2 I'm most happy with: Agent-first operation. The harness ships its own manual as auto-triggered skills, so you drive it by just talking to a Claude Code or Codex session — it loads the right skill and runs the commands for you. Agents driving agents: zero flags to memorize, zero learning curve. Dual-SDK support. Every task can run on Claude Code or OpenAI Codex, selectable per agent. Switch seamlessly, or split work across both — may the best model win. It's source-available (free for noncommercial use). Code is at https://lnkd.in/gH6SHhwy The v2 harness has let me run 3-5 projects concurrently — easily maxing out a top-tier subscription quota — including developing the harness itself. Curious what harness or loop you're running now — what do you like and dislike about it? #AgenticCoding #SoftwareEngineering #AIEngineering #LLM
-
Yuntao Jia liked thisWild idea turns into world changing brand; hardwork result in today’s IPO. Thank you Limers, for all the innovation and hardwork over years, through touch times and thrive from Pademic! This is Day One, this is time to build. Ride on!Yuntao Jia liked thisAs the world's largest shared electric vehicle company, Lime has a mission to build a future where transportation is shared, affordable, and carbon-free. Here's to a new chapter, $LIME! #NasdaqListed
-
Yuntao Jia liked thisThank you guys!!!Yuntao Jia liked thisRoughly 10 years ago, I cofounded Lime together with Toby Sun and Adam Zhang, with a dream to #Unlock_Life by providing users a new way that's convenient, affordable, clean and safe way to navigating the city like never before. As founder, I’m proud of how far we’ve come—from an idea to a global platform helping reshape urban mobility. Still remember the day we celebrated 1M rides, now we are at 1 Billion Rides. With the company continuing to scale and mature, the time feels right for a transition in board leadership, so I can dedicate more of my time to building my next venture - Pear.Us, to #Unlock_Live. Pear is the Commerce Layer for Culture, Pear aim to bring the most authentic experience to fans, and attribute the data and economics back to who creates them. I’m excited to share that Jim Rowan will be stepping into the role of Chairman of the Board. Jim brings deep operational experience, global perspective, and strong leadership that will help guide Lime through its next chapter of growth. I will remain actively involved as a member of the board, continuing to support Wayne Ting, the team and the mission we started. Grateful to our team, partners, riders, and investors who have been part of this journey—and excited for what’s ahead. — Brad Bao
-
Yuntao Jia liked thisAt our second annual TikTok Shop U.S. Summit this April in Los Angeles, we brought together brands, sellers, and partners to explore how commerce is being reshaped through discovery. Key takeaways: - Demand is increasingly driven by content, not just intent - Creator and LIVE ecosystems are accelerating conversion - Discovery is enabling more sustainable, repeatable growth Read more in Fortune on how discovery is driving durable commerce 👇 #TikTokShop #DiscoveryCommerce #TikTokShopSummit2026Yuntao Jia liked thisIn the global discovery e-commerce marketplace, brands are learning how to turn attention into something that lasts. Here’s how companies are harnessing the power of TikTok Shop. Learn more: https://lnkd.in/emJsRVCnHow TikTok Shop is turning discovery into durable commerce | FortuneHow TikTok Shop is turning discovery into durable commerce | Fortune
-
Yuntao Jia liked thisYuntao Jia liked thisI'm excited to announce I joined as CTO of Runbook (and I'm hiring!). We're building the AI workforce for the physical economy — AI agents that replace the manual coordination layer running on emails, spreadsheets, TMS, ERP, and phone calls at companies that move, make, and deliver physical things. Our agents handle complex multi-day workflows, learn from feedback, and reach 90%+ autonomous execution in production. The engineering problems are real, the team is small and mighty. We have multiple customers including a Fortune 500. We move fast and hold a high bar. I'm looking for founding engineers who've shipped at scale and want real ownership over something hard!
-
Yuntao Jia liked thisYuntao Jia liked thisDuring the Super Bowl Breakfast last Saturday, LiveX AI was honored to feature a fully interactive hologram of Christian McCaffrey as he was celebrated for receiving the 2026 Bart Starr Award. This marked the first time an AI athlete engaged directly with fans while amplifying the cause that matters most to him. Christian was the right athlete to lead that moment. Through his performance on the field, his connection with fans, and his support of military personnel and their families, his leadership reaches far beyond football. Watching fans share in that recognition through the interactive hologram was one of the highlights of the weekend, and we were thrilled to capture some of those reactions in the video below. For LiveX and the sports world, this milestone also showed how AI can help athletes extend their message, deepen fan connection, and carry their impact beyond the game. We want to thank Armada for bringing us into the event and for their work helping push America’s AI capabilities forward. Like LiveX AI, Armada sees how AI can enhance the fan experience at live events and believes in a future where events are more interactive and memorable than ever before. We had great conversations with the Armada team about what’s possible when our technologies come together: LiveX powers the AI-to-fan interaction layer, while Armada could provide the mobile data center infrastructure that makes real-time experiences like this possible anywhere. And finally, thank you Marriott Marquis San Francisco and Athletes in Action for putting on such an impactful event. The breakfast was a peak moment of the week and set the tone for an unforgettable Super Bowl weekend.
-
Yuntao Jia liked thisYuntao Jia liked thisToday marks the official close of our acquisition by ServiceNow. Thank you to Amit Zavery, Bill McDermott, and the ServiceNow team for the trust and partnership as we take this next step together. We started in what I called the Machine Learning era, when building a production-grade question-answering system felt almost impossibly hard. In the first half of the 9 year journey, we built a top-of-class AI assistant for work with the technology of that time, pushing far beyond what was considered possible back then. It worked. But then the ground shifted. The move to Agentic AI was a technical reset for us. Pivoting at that stage wasn’t easy, especially with hundreds of customers relying on the system every day. We had to rebuild the plane while flying it: start over architecturally, innovate against time, and make hard calls fast. Through it all, the real story was the people. We stuck together, we kept building, and we prevailed. As a first-time founder, I’m deeply grateful to my co-founders, Bhavin Shah, Vaibhav N., and Varun Singh, for showing me the way when I clearly didn’t know what I was doing. And to our team: Chang Liu, Yi Liu, Jing Chen, Manan Gosalia, Sadish Ravi, Gwendolyn Thorn, Ivy Wang, Dave Uppal, Desmond Chan, Matt Mistele, Cody Kala, Jinghui Yu, Lucas Liu, Karey Shi, Tony Jin, Ajay Merchia, Ritwik Raj, Andrew MairenaKerman Lau, AnnMarie Zimmermann, Adam Levy, Cecilia (Cici) Cao and so many others, thank you for believing, staying, and pushing through. For nine years, you didn’t just build a company; you built a legacy. With the strength of ServiceNow's platform and reach, we are more uniquely positioned than ever to take on the challenge of agentizing the enterprise, every workflow, every API, every piece of data. Onward! https://lnkd.in/gK6KgTu8
-
Yuntao Jia liked thisExcited to join Ishan Gupta 🧃, David Paffenholz 🧃 and the amazing team at Juicebox in building the future of recruiting! We're hiring across a variety roles - come build with us: http://juicebox.ai/careersYuntao Jia liked thisWe’re excited to welcome Mack Yi to Juicebox as our new Engineering Lead. What stood out immediately was his range of past experiences. He’s spent the last few years at Granica leading engineering for new products and taking ideas from zero to production. Before that he spent five years at Airbnb working on reliability, categorization, and large scale data pipelines across some of their highest traffic surfaces. It’s rare to find someone who’s equally comfortable in the early stage chaos and in environments where every system has to scale to millions of users. Mack’s only been here a short time but he’s already deep in one of our most requested feature areas. The full details are coming soon, but it will help companies put their existing data to work and surface people they should be talking to much faster. He’s going to be a big unlock for what we’re building next, and we’re thrilled he’s a part of the team!
Experience & Education
-
One Market
********** * ***
-
******
**** ** ***********
-
******
*********** *******
-
********** ** ******** ****************
*** ******** ******* undefined
-
-
******** **********
****** ******** *******
-
View Yuntao’s full experience
See their title, tenure and more.
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
or
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Languages
-
English
Native or bilingual proficiency
-
Chinese
Native or bilingual proficiency
Recommendations received
3 people have recommended Yuntao
Join now to viewView Yuntao’s full profile
-
See who you know in common
-
Get introduced
-
Contact Yuntao directly
Other similar profiles
-
Sumanth Kolar
Sumanth Kolar
Experienced tech executive with a track record of scaling fast growing products and teams. Skilled at building high-performing engineering teams, strategic thinking and collaborating effectively with cross-functional executives. Currently on a break - advising startups and traveling. Previously, Head of Engineering for Uber Rides team (300+ eng, Uber app, backends, AI & Mobility verticals), Sr Director at LinkedIn (News Feed, Content & Jobs), Head of Engineering at SlideShare
4K followersPalo Alto, CA
Explore more posts
-
Amir Zohrenejad
Heavybit • 2K followers
Post-training is now driving most improvements in foundation models, and billions are being poured into RL. But while RL might look like the next obvious infrastructure category, I’m not convinced it is. Inference scaled because it was an engineering problem. RL and post-training are different. The hard parts (environments, reward design, and evaluation) require domain expertise in the task and trial and error. A viable post-training infra company likely needs to offer something meaningfully new: continual, on-policy learning as a service. The service would need to provide: 1. Pipelines that synthesize RL environments directly from model outputs 2. Systems that construct (or at least propose) reward signals and verification 3. Background updates to model weights with minimal human intervention Full post in the comments.
19
2 Comments -
Johanan Ottensooser
fiveonefour • 5K followers
Week 1 of #OlapOnTap covered: 1. Typing: how much it can matter in OLAP contexts: https://lnkd.in/gCQnNTNt 2. A deep dive into Cardinality, how it affects performance, and how using enums and LowCardinality() can help improve your performance: https://lnkd.in/g2d_pQJZ 3. Informed by this, how to efficiently design your table's ORDER BY to maximize your database's performance: https://lnkd.in/g2BnBCFF 4. How denormalization, and wider, flatter tables, are a better fit for OLAP database engines and query patterns https://lnkd.in/gyQQ62Gr The TL;DR for Week 1 is: 🤏 More restrictive types drastically improve performance, and low cardinality types can improve that even more 🔭 Structure your data based on your consumption patterns (both in terms of how you order your data, and how you structure your tables). Next topics coming up are: ➜ Modeling JSON & semi-structured data (noting how much heavy lifting the word "semi" is doing in this title) ➜ Nullability and default values ➜ Compression and encoding (and how you can take the reins more than you think here) ➜ Precomputing with materialized views The starting theme is modeling / setting up your OLAP database (if that wasn't obvious!). I am just kinda exploring topics I find interesting in the OLAP space, and I want to write LLM documents for my agents in respect of. But if any of you have particular topics you want me to cover, holler! Also, I'm focusing on ClickHouse implementations, but let me know if you want me to cover other databases. #DataEngineering #OLAP #MooseStack
11
-
Purusottam Mupunu
Cloudanix • 6K followers
LLMs are amazing at generating code. However, one of the biggest limitation of this code isn’t intelligence - it’s lack of context. Things get interesting (and messy) when it comes to large codebases. It falls apart and it's painfully obvious why. LLMs can generate working implementations from a high-quality prompt. At their best, they can produce a correct feature in one shot (“one-shotting”). This helps dramatically speed up development for clearly scoped tasks. This helps with: ✅ Productivity boost for implementation-heavy tasks - For repetitive CRUD operations or isolated scripts, LLMs can collapse hours of work to minutes. ✅ Enhanced prototyping — They help quickly explore solutions, enabling engineers to iterate design ideas faster. ✅ Lower entry barrier — Less context switching (e.g., from editor to docs), faster feedback loops. But these benefits fade when tasks require deep architectural knowledge, modular context, or long dependency chains, especially in large codebases. In this blog post Kieran Gill from Blueberry Pediatrics highlights some of the reasons it fails and how it can be improved. Here are a few reasons why LLMs fail with large codebases: ❌ Lack of global context: LLMs don’t “understand” the whole system - they only see what’s in the prompt. ❌ Hallucinations in unfamiliar parts: When missing architectural knowledge, LLMs may generate incorrect code, leading to bugs or regressions. ❌ Rework becomes more expensive than doing it manually: Failing to one-shot means repeated rounds — which erodes the time savings. Here's how you can solve the challenges: ▶️ Guide LLM using Prompt libraries that show Architecture, best practices, domain knowledge, etc. ▶️ Improve the codebase by making code modular, consistency in naming, and keep code clean. ▶️ Invest in review & automation for verification of LLM's design choices and end product (code, unit tests, etc.) In large codebases, the real productivity is gained by guidance + oversight + modular architecture. #ai #llm #aiengineering #engineering https://lnkd.in/gFzFunEV
18
-
Jimin Lee
Samsung Electronics • 2K followers
Everyone says LLMs are getting better. True. But shipping an LLM that users actually love still requires significant craft, meticulous engineering, and a focus that goes far beyond the model itself. The gap between a "cool demo" and a "reliable product" is where the real work happens. Here are the key areas that demand attention for building high-quality, production-ready LLM applications: 1. Prompting still matters. Models have improved, but they can’t read your mind. A well-crafted prompt can make results shine, while a poor one can make them fall flat. 2. Design for prompt agility. Build a deployment path where you can update prompts/configs without a full release. (Feature flags, versioned prompts, audit trails.) 3. Language mixing is real and sneaky. Sometimes the LLM itself slips in the wrong language mid-sentence. Without careful handling, your output ends up a confusing mix instead of what the user expects. 4. Multilingual isn’t free. Don’t assume the LLM will magically handle every language. You still need to evaluate results carefully and when things break, be ready to fix them with prompt engineering or fine-tuning. 5. 80% quality is “demo-easy.” 95% is war. The last 15% is where you pay: edge cases, latency tails, safety filters, caching strategy, and human UX. 6. Evaluation is annoying but non-negotiable. Do both: - Small, focused human evals for truth, nuance, and unexpected failure modes. - Large, automated runs (LLM judges + task metrics) for coverage and speed. Ship a repeatable eval pipeline before you ship features. 7. Integrate with the actual app early. LLMs don’t live in notebooks. Think auth, rate limits, retries, drafts/confirm flows, offline states, and analytics from day one. 8. Write failures into the UX. Hallucinations, timeouts, and “I don’t know” need graceful fallbacks. 9. Latency is a product feature. Users forgive a small mistake; they don’t forgive lag. Token streaming, caching, and truncation matter. 10. Guardrails are not just one filter. Stack them: prompt hardening, retrieval constraints, post-filters, and red-team tests. Defense in depth. 11. Training is expensive. Inference is more expensive. Training burns a hole in your budget once. Inference burns it forever - every request, every user, every day. Optimize throughput and cost-per-request early. Treat prompts like code, eval like CI, and the LLM like one component in a system not the whole system.
18
-
Juan Destribats
Listo • 1K followers
Everyone is debating the latest model releases and benchmark scores. But the real competition isn’t at the LLM layer anymore, it’s shifting to the application layer that sits on top of chat interfaces. Models will keep improving, but they’re converging. The moat is shifting from raw model intelligence to context and memory. Platforms that accumulate user intent, preferences, and behavioural context fastest will hit escape velocity. That’s why ChatGPT Apps matter so much. 🚀 They create a value exchange: third-party developers get organic distribution, and the platform gains more usage, more intent signals, and more personalized context to improve the experience for every app More apps → more usage → more user context → better app experiences → more apps. And round it goes. A full flywheel. ⚙️ But here’s the subtlety people miss: Apps and products become more powerful in this model, not less. They plug directly into an interface that already understands the user. Authentication, preference history, behavioural patterns, and previous queries all become inputs to your app’s experience. In every platform shift, visibility goes to the early movers. AI will be no different. Build your AI-app with Listo
43
2 Comments -
Roxane Fischer
Anyshift.io • 8K followers
On Tuesday we hosted a live product demo for Anyshift in San Francisco! 🔥 Thanks to Hilary Brennan-Marquez from MotherDuck for being part of it! Two new features shown for the first time: - Annie can now generate architecture diagrams on the fly from the full infrastructure graph - Proactive Annie continuously monitors your production, surfacing weak signals and monitoring gaps before they become incidents Hilary joined us from New York 🏙️ to share how her team uses Annie at MotherDuck. With a team that tripled in two years, Annie became the go-to partner for junior engineers during on-call and onboarding. The team runs everything through Slack and now has an "Ask Annie" channel where even non-infra engineers drop in with questions. More demos coming soon 👊 Reach out if you want to see Annie in action.
16
-
Paul-Anthony Dudzinski
Amazon • 1K followers
I read a fascinating paper over the weekend that continues to convince me neuroscience and AI model research are eventually related tracks: https://lnkd.in/eTGYtG4x. Researchers discovered they can identify specific neurons in LLMs that predict when the model is about to hallucinate. Even more striking: less than 0.1% of neurons are responsible for these errors. Think about that. Just like neuroscientists can pinpoint which parts of the human brain light up during specific activities, we can now see which AI neurons are firing when a model generates false information. The team traced these neurons back to pre-training, showing they emerge early in the model's development. This is very similar to how certain neural pathways form in our own brains. The research reveals these neurons are causally linked to over-compliance behaviors, where models prioritize generating plausible-sounding responses over admitting uncertainty. What excites me most is the practical potential. We can develop better training approaches and even real-time detection systems. The researchers expanded on previous work to identify multiple categories of hallucination causes, giving us a more nuanced view of why models fail. Really looking forward to how the research evolves on this one!
12
1 Comment -
Julia Proskurnia
Google • 3K followers
Weekly arXiv Scan: 3 Papers that caught my eye (and my Agent’s) My custom Gemini x OpenClaw agent scanned recent papers on arXiv this week in Audio, NLP, and GenAI. It filtered for novelty and engineering rigor. I filtered for practical application and "human" impact. Here are the winners that made the cut during the baby nap time :) 1. Scaling Open Discrete Audio Foundation Models (SODA) https://lnkd.in/eXYnjfpd The Agent's TL;DR: Introduces SODA, a 4B parameter native audio model trained on interleaved semantic, acoustic, and text tokens, establishing new scaling laws for non-cascaded speech tasks. My Take: Human conversation relies on tone, pauses, and emotion. Those are the things that get stripped away when we convert Speech-to-Text for an LLM. By processing audio tokens natively, we aren't just processing data; we are preserving the humanity of the interaction. Native audio-in architecture is what we need for empathetic AI. 2. Reverso: Efficient Time Series Foundation Models https://lnkd.in/eyAgsN_h The Agent's TL;DR: Proposes a hybrid architecture interleaving linear RNNs and long convolutions that matches Transformer zero-shot forecasting performance while being 100x smaller in parameter count. My Take: Time-series analysis and prediction might be the problem where the sledgehammer solution is an overkill. A model that is 100x smaller is a model one can actually use for the personal gains (like some extra signals for investments? maybe?). It’s faster to train, cheaper to serve, and easier to debug. 3. A Generative-First Neural Audio Autoencoder https://lnkd.in/eeFtDpkB The Agent's TL;DR: Proposes an architecture with a massive 3360x temporal downsampling factor, allowing 60 seconds of audio to be represented by fewer than 800 tokens for extremely fast generation. My Take: High fidelity is great, but latency is the user experience killer. Compressing a minute of audio into <800 tokens is a massive engineering win. This reduces the context burden on the LLM significantly, meaning faster, cheaper, and more responsive voice agents. This is a great step towards the feeling of a snappy real-time conversations. The trend this week is “Efficiency over Excess”, finding architectures that do the same work with a fraction of the compute. #NLP #AudioAI #MachineLearning #HumanInTheLoop #arXivScan
41
8 Comments -
Jayan Tharayil
Pinch AI • 2K followers
While I’m over here crafting "𝙘𝙖𝙩𝙘𝙝𝙮" LinkedIn hooks and fluffy strategy pieces, the actual geniuses at Pinch AI are busy engineering 𝟵𝟬% 𝗽𝗲𝗿𝗳𝗼𝗿𝗺𝗮𝗻𝗰𝗲 𝗴𝗮𝗶𝗻𝘀. They just moved our feature runs from hours to minutes. 🚀🚀 🦀 I’ll stick to the emojis. 🦀 If you want to see what real technical heavy lifting looks like (and why I leave the data architecture to people far smarter than me), check out Mandy's (Mandeep Kaur) breakdown on our switch to DuckDB.
12
-
ROMAN SHRAMKOV
SOKILDEV • 5K followers
80.3% on SWE-Bench Pro. Anthropic held that number back from public release for three months, judging it too disruptive. As of June 9, any team with an API key can run it. Claude Fable 5 is the first Mythos-class model opened to general developers. The key capability: full codebase-wide migrations completed in a single day. If you've been estimating migration work in weeks, that math has changed. Worth knowing before you ship anything with it. Anthropic routes sensitive queries - biosecurity, cyberweapons - to a weaker model rather than refusing outright. The company retains all traffic data for 30 days under its own security policy. The higher-capability Mythos 5 tier went to existing Project Glasswing partners at the same time - public Fable 5 is not the ceiling. Separately, the White House executive order signed June 2 requires a mandatory 30-day pre-release government testing window for AI models going forward. If you're building on this stack long-term, that review layer matters for release timeline planning. At SOKIL we run Java/React delivery teams with AI-augmented workflows. Honest read: Fable 5 is the first model where we'd scope a migration task as an agent task, not a human task with AI assistance. That changes how we write estimates and how we talk to clients about timelines. If your competitors start running autonomous agents on migration work, when does that start showing up in their proposals to your customers? #AI #EngineeringLeadership #DeveloperTools #Anthropic #LLM
3
-
Jacob Clark
Hyperact • 13K followers
Why do LLMs return different responses to the same prompt? 🤔👉 Interestingly enough the Transformer architecture underpinning Large Language Models (LLMs) is not inherently non-deterministic, given identical inputs it is capable of computing identical outputs! So why does the same prompt sometimes produce a different response in systems like ChatGPT or Claude? Non-determinism enters these system through two sources: design decisions and operational entropy. 1️⃣ By Design These systems don’t just "pick an answer", they generate text by sampling from a probability distribution of possible next tokens (small sequences of characters) and continually feed each selected token back through the system. So instead of always choosing the most probable token (known as greedy decoding), they intentionally sample to produce more natural human-like text. Factors such as: - Temperature settings control how varied the system can be when selecting the next token, higher temperature values yield more creative outputs - Top-k sampling limits reshape the probability space before selection which also constrains which can also help to increase or reduce perceived determinism balanced with diversity of responses We have to remember that these systems are designed to achieve their objectives as a human might, this design choice introduces intentional diversity in responses, non-determinism here is not a flaw. 2️⃣ Operational Entropy Even with temperature set to zero and a very small top-k, responses can still differ. That’s because floating-point arithmetic on GPUs and TPUs is non-associative, meaning the order of operations matters when the math is performed. When thousands of requests are processed concurrently, tiny timing or batching differences cause slight numerical shifts in the systems logits. These small deviations cascade producing subtle variations in all to be generated tokens. This of course compounds in very large responses. This is not due to inherent randomness in the base models themselves, but to the realities of distributed, parallel computation in a highly stateful system of this scale. 😬 So, non-determinism is both a feature and a by-product: - By design, it enables creativity and linguistic diversity - By circumstance, it reflects the unavoidable quirks of floating-point math and at scale inference Even “deterministic” configurations can still vary slightly which is why two identical prompts don’t always yield identical answers. (I’ve intentionally used the term "system" rather than "model" to reflect how modern tools like ChatGPT and Claude are no longer isolated models deployed within an inference stack, but complex, orchestrated systems in their own right, combining multiple models, retrieval mechanisms, control layers, and interaction policies to deliver coherent, adaptive behaviour.)
17
3 Comments
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content