Skip to content
View kekmodel's full-sized avatar

Block or report kekmodel

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. THUDM/slime THUDM/slime Public

    slime is an LLM post-training framework for RL Scaling.

    Python 8.4k 1.2k

  2. FixMatch-pytorch FixMatch-pytorch Public

    Unofficial PyTorch implementation of "FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence"

    Python 798 185

  3. MPL-pytorch MPL-pytorch Public

    Unofficial PyTorch implementation of "Meta Pseudo Labels"

    Python 389 64

  4. rl_pytorch rl_pytorch Public

    Deep Reinforcement Learning Algorithms Implementation in PyTorch

    Jupyter Notebook 27 4

  5. reinforcement-learning-kr/alpha_omok reinforcement-learning-kr/alpha_omok Public

    Minimal version of DeepMind AlphaZero

    Python 84 22

  6. reinforcement-learning-kr/distributional_rl reinforcement-learning-kr/distributional_rl Public

    Repository for studying distributional rl

    Python 30 8