Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
reinforcementlearning
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
How Early Digital Systems Quietly Shaped the Minds Building Tomorrow
Alan Scott Encinas
Alan Scott Encinas
Alan Scott Encinas
Follow
Jul 28
How Early Digital Systems Quietly Shaped the Minds Building Tomorrow
#
autonomousvehicles
#
reinforcementlearning
#
simulationtraining
Comments
Add Comment
5 min read
Decoding the Link Between Pretraining and Reinforcement Learning
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 20
Decoding the Link Between Pretraining and Reinforcement Learning
#
airesearch
#
largelanguagemodels
#
reinforcementlearning
#
machinelearning
Comments
Add Comment
3 min read
Muon Optimizer Boosts Agentic Reinforcement Learning Performance
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 20
Muon Optimizer Boosts Agentic Reinforcement Learning Performance
#
reinforcementlearning
#
muon
#
optimizer
#
agenticai
Comments
Add Comment
3 min read
The ~+9.4% You Can't Afford to Verify: Evaluating SDAR (and the FinOps of Trying)
Shoaibali Mir
Shoaibali Mir
Shoaibali Mir
Follow
Jul 18
The ~+9.4% You Can't Afford to Verify: Evaluating SDAR (and the FinOps of Trying)
#
aws
#
machinelearning
#
mlops
#
reinforcementlearning
Comments
Add Comment
6 min read
Length Penalties in LLMs: Shorter Chains of Thought, Hidden Influences
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 17
Length Penalties in LLMs: Shorter Chains of Thought, Hidden Influences
#
llm
#
chainofthought
#
reinforcementlearning
#
interpretability
Comments
Add Comment
3 min read
Agent Apprenticeship turns finished agent tasks into reusable experience
Reno Lu
Reno Lu
Reno Lu
Follow
Jul 16
Agent Apprenticeship turns finished agent tasks into reusable experience
#
aiagents
#
posttraining
#
agenttraces
#
reinforcementlearning
Comments
Add Comment
3 min read
I Replaced a Q-Table With a Neural Network and Everything Changed - Day 5 (DQN).
Madhumitha Kolkar
Madhumitha Kolkar
Madhumitha Kolkar
Follow
Jul 24
I Replaced a Q-Table With a Neural Network and Everything Changed - Day 5 (DQN).
#
reinforcementlearning
#
machinelearning
#
deeplearning
#
python
5
 reactions
Comments
Add Comment
4 min read
AI Agents Are Learning to Build the Worlds They Train In
Breach Protocol
Breach Protocol
Breach Protocol
Follow
Jul 1
AI Agents Are Learning to Build the Worlds They Train In
#
aiagents
#
worldmodels
#
reinforcementlearning
#
alibaba
Comments
1
 comment
4 min read
Why teaching AI agents to use tools keeps blowing up in training
Breach Protocol
Breach Protocol
Breach Protocol
Follow
Jul 2
Why teaching AI agents to use tools keeps blowing up in training
#
reinforcementlearning
#
agents
#
tooluse
#
training
Comments
Add Comment
3 min read
Building a Self-Optimizing Python Trading Bot with Reinforcement Learning and Binance API
Fazil Hasanov
Fazil Hasanov
Fazil Hasanov
Follow
Jun 19
Building a Self-Optimizing Python Trading Bot with Reinforcement Learning and Binance API
#
python
#
trading
#
reinforcementlearning
#
binance
Comments
Add Comment
4 min read
The Whole Paper Fits in One Sigmoid: Implementing the SDAR Gate
Shoaibali Mir
Shoaibali Mir
Shoaibali Mir
Follow
Jun 14
The Whole Paper Fits in One Sigmoid: Implementing the SDAR Gate
#
machinelearning
#
reinforcementlearning
#
python
#
aws
Comments
1
 comment
5 min read
Four Models in One Training Loop: Architecting SDAR on AWS (Before Renting a Single GPU)
Shoaibali Mir
Shoaibali Mir
Shoaibali Mir
Follow
Jun 6
Four Models in One Training Loop: Architecting SDAR on AWS (Before Renting a Single GPU)
#
aws
#
machinelearning
#
reinforcementlearning
#
mlops
Comments
Add Comment
5 min read
How to Add Live Telemetry and Failure Diagnosis to Isaac Lab, MuJoCo, or Gazebo Training in Under 5 Minutes
SimTooReal
SimTooReal
SimTooReal
Follow
Jun 6
How to Add Live Telemetry and Failure Diagnosis to Isaac Lab, MuJoCo, or Gazebo Training in Under 5 Minutes
#
ai
#
robotics
#
mujoco
#
reinforcementlearning
Comments
Add Comment
4 min read
Why robotics RL training pipelines fail at scale
Robosynx
Robosynx
Robosynx
Follow
May 30
Why robotics RL training pipelines fail at scale
#
robotics
#
machinelearning
#
reinforcementlearning
#
simulation
Comments
Add Comment
4 min read
ARTIST: RL-Powered Tool Use for LLM Agents Explained
Jangwook Kim
Jangwook Kim
Jangwook Kim
Follow
May 27
ARTIST: RL-Powered Tool Use for LLM Agents Explained
#
reinforcementlearning
#
llmagents
#
tooluse
#
agenticai
Comments
Add Comment
9 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account