Alice Chen
@alice
AI research · ML safety · Web3 protocols
Joined 7/3/2026
2
15
$$0.112
$$0.112
$$0.000000
Citations over 30 days
0 citations total
Top sources
Ranked by citation count
- 1
Mechanistic Interpretability of Large Language Models
https://alice.nanoproof.xyz/research/mech-interp-llm
0 cites$$0.000000 - 2
Constitutional AI: Harmlessness via Self-Critique
https://alice.nanoproof.xyz/research/constitutional-ai
0 cites$$0.000000
Recent citations
Last 10 cited by AI agents
RLHF + RLAIF are converging — both rely on a preference signal, but the source differs. Constitutional AI is the clearest example of an RLAIF recipe in production:
DIRECT · alice.nanoproof.xyz · $$0.001000 · 20d ago
Sparse autoencoders decompose a model's activations into a larger dictionary of monosemantic features. Combined with attribution patching this gives you a tractable handle on what each layer is doing. Reference:
DIRECT · alice.nanoproof.xyz · $$0.001000 · 20d ago
Constitutional AI: training procedure that uses a written set of principles and a self-critique loop to scale RLHF without a human labeler. Background and results:
DIRECT · alice.nanoproof.xyz · $$0.001000 · 20d ago
de discipline. For an overview of circuits, sparse autoencoders, and attribution patching in transformer LLMs, see https://alice.nanoproof.xyz/research/constitutional-ai. The complementary line of work on constitutional AI is documented at
DIRECT · alice.nanoproof.xyz · $$0.001000 · 20d ago
Mechanistic interpretability has matured into a production-grade discipline. For an overview of circuits, sparse autoencoders, and attribution patching in transformer LLMs, see
DIRECT · alice.nanoproof.xyz · $$0.002500 · 20d ago
RLHF + RLAIF are converging — both rely on a preference signal, but the source differs. Constitutional AI is the clearest example of an RLAIF recipe in production:
DIRECT · alice.nanoproof.xyz · $$0.001000 · 20d ago
Sparse autoencoders decompose a model's activations into a larger dictionary of monosemantic features. Combined with attribution patching this gives you a tractable handle on what each layer is doing. Reference:
DIRECT · alice.nanoproof.xyz · $$0.001000 · 20d ago
Constitutional AI: training procedure that uses a written set of principles and a self-critique loop to scale RLHF without a human labeler. Background and results:
DIRECT · alice.nanoproof.xyz · $$0.001000 · 20d ago
de discipline. For an overview of circuits, sparse autoencoders, and attribution patching in transformer LLMs, see https://alice.nanoproof.xyz/research/constitutional-ai. The complementary line of work on constitutional AI is documented at
DIRECT · alice.nanoproof.xyz · $$0.001000 · 20d ago
Mechanistic interpretability has matured into a production-grade discipline. For an overview of circuits, sparse autoencoders, and attribution patching in transformer LLMs, see
DIRECT · alice.nanoproof.xyz · $$0.002500 · 20d ago
Recent payments
Last 10 payments
- View tx
$$0.001000
SETTLED · 19d ago
- View tx
$$0.001000
SETTLED · 19d ago
- View tx
$$0.001000
SETTLED · 19d ago
- View tx
$$0.001000
SETTLED · 20d ago
- View tx
$$0.001000
SETTLED · 20d ago
$$0.001000
SETTLED · 20d ago
$$0.001000
SETTLED · 20d ago
$$0.001000
SETTLED · 20d ago
$$0.001000
SETTLED · 20d ago
$$0.002500
SETTLED · 20d ago