I turn capability gaps into trainable, verifiable systems.

I am an Applied Scientist at Amazon AGI, working on frontier-model post-training. My work spans reasoning environments, verifiers, curriculum, rollout systems, reinforcement learning, and evaluation.

Yuhua (Bill) Chen

The difficult part is not producing more tasks or more rollouts. It is keeping the training signal aligned with the intended capability as the policy, environment, and evaluator all change.

01

Define the capability

A benchmark gap is not automatically the capability gap.

02

Make it trainable

A difficult task is not automatically a useful learning task.

03

Verify the behavior

A correct checker is not automatically a trustworthy reward.

04

Adapt the curriculum

A fixed distribution does not stay informative as the policy changes.

05

Run the system

Reliability failures can change the experiment, not merely delay it.

06

Test what transferred

A benchmark improvement is not automatically capability acquisition.

Evaluation is useful when it changes the next formulation: what to train, what to verify, and which apparent gains deserve trust.

READ THE TECHNICAL VIEW →

Before frontier models, the closed loop was an autonomous MRI scanner.

At Q.bio, I deployed 3D AI systems inside an autonomous MRI scanner, including components with sub-second inference for real-time localization and imaging decisions.

My earlier work in 3D MRI reconstruction, super-resolution, and segmentation has been cited more than 1,000 times.

I am a co-inventor on patent applications covering deep-learning MRI reconstruction and calcium-free CT angiography.

RESEARCH & TECHNICAL WORK →
Across medical imaging and frontier-model post-training, my work has followed the same pattern: identify where the current formulation breaks down, redefine the problem around the real bottleneck, and build the full system needed to test the new formulation in practice.
ABOUT THE ARC →

Interested in hard problems around reasoning and agent capabilities?