HuArm
Collaborative Robotic Erhu Bowing via Hierarchical Reinforcement Learning



Project Overview & Description
Robotic musicianship is an active area of research, enabling new methods of instrument interfacing and sound production. However, expressive bowing for bowed-string instruments has remained unsolved. As an erhu player myself, I chose to focus on this Chinese folk instrument, in the hope that the results can generalize to other bowed-string instruments.
HuArm is a custom 5-DoF robotic arm trained in a JAX-based MuJoCo (MJX) pipeline using hierarchical reinforcement learning, with the goal of safe, adaptive collaborative music-making alongside human players. Solver and contact complexity are tuned against physics realism to speed up policy learning, and domain randomization over sensor noise and physics parameters makes the policy robust for sim-to-real transfer. Heuristics for proper erhu bowing — such as keeping the bow flat and close to the sound box — are embedded directly into the environment's reward function. Trajectories are generated to model realistic bowing motions, with varying length, velocity, and pressure.
Currently, this project achieves successful and safe tracking of bow velocity and pressure in simulation, across arbitrary trajectories. Beyond the technical contribution, HuArm has potential as a conduit for cultural preservation and music education, and points toward an exciting future application: an EEG-conditioned prosthesis for disabled erhu players.
Project contributions
- A hierarchical reinforcement learning policy for safe, adaptive bowing alongside human players.
- A JAX-based MuJoCo pipeline with an erhu-specific string-contact model, tuned for training speed against physics realism.
- Domain randomization over sensor noise and physics parameters for a sim-to-real robust policy.
- An iPhone teleoperation pipeline that collects dexterous bowing contact data for imitation learning and policy warm-starting.