Fremont, California, United States
8K followers 500+ connections

Join to view profile

About

Technologist at heart. Always learning, always building — compilers, distributed systems,…

Articles by Yun

  • Why Benchmark Parity Wasn't Enough

    In The Same Weights Are Not the Same Model, I argued that an inference provider owes users fidelity: serving the model…

    2 Comments
  • The Rack Is the New Node: NVL-72 and Frontier Inference

    If you have spent the last twenty years building internet services, systems like GB300 NVL-72 might feel wrong —…

    19 Comments
  • The Same Weights Are Not the Same Model

    Earlier this month Dmytro Dzhulgakov , one of Fireworks AI's co-founders, posted a sentence resonating with me a lot:…

    1 Comment
  • Kimi K3 Cost Matrix

    Kimi K3 was released last week. Fireworks AI enabled it on day0, so did the open source engines vLLM and SGLang.

    1 Comment
  • The Best AI Isn't One Model — It's a Choice

    The math and the money behind routing across many models instead of betting on one. For two years, "use the best AI"…

    4 Comments
  • Responsible tokenmaxxing with open-weight models

    "Tokenmaxxing" - pointing AI at as much valuable engineering work as possible and letting it complete more of it…

    3 Comments
  • To Own, or to Rent: That is the AI Question

    My boss, Lin Qiao,wrote this week about owning vs. renting intelligence — and she's right: the conversation has been…

    5 Comments
  • Attention Is All You Need. Not All Attention Is Needed.

    Language is far sparser than the original Transformer assumed. The last 18 months of model architecture are the…

    1 Comment
  • Inference Is No Longer Cost per Token. It Is Cost per Task.

    There's a line Jensen Huang dropped when asked how much inference compute would grow: "It's about to go up a billion…

    7 Comments
  • The Myth of "Working Less": managing AI is the new job

    We've all seen the headlines: AI will automate the busywork, clear the calendar, and finally give us back our time—as…

    4 Comments

Activity

8K followers

See all activities

Experience & Education

  • Fireworks AI

View Yun’s full experience

By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.

Publications

Patents

  • Assigning type parameters

    Issued US 20110302555

    Other inventors
    See patent
  • Scheduling by Growing and Shrinking Resource Allocation

    Filed EU EP20080826472

    A scheduler for computing resources may periodically analyze running jobs to determine if additional resources may be allocated to the job to help the job finish quicker and may also check if a minimum amount of resources is available to start a waiting job. A job may consist of many tasks that may be defined with parallel or serial relationships between the tasks. At various points during execution, the resource allocation of active jobs may be adjusted to add or remove resources in response…

    A scheduler for computing resources may periodically analyze running jobs to determine if additional resources may be allocated to the job to help the job finish quicker and may also check if a minimum amount of resources is available to start a waiting job. A job may consist of many tasks that may be defined with parallel or serial relationships between the tasks. At various points during execution, the resource allocation of active jobs may be adjusted to add or remove resources in response to a priority system. A job may be started with a minimum amount of resources and the resources may be increased and decreased over the life of the job.

    Other inventors
    See patent

Languages

  • English

    -

  • Chinese

    -

View Yun’s full profile

  • See who you know in common
  • Get introduced
  • Contact Yun directly
Join to view full profile

Other similar profiles

Explore top content on LinkedIn

Find curated posts and insights for relevant topics all in one place.

View top content

Add new skills with these courses