~/vijay
Photo of Vijay Murari Tiyyala

vijay@bu:~$ whoami

Vijay Murari Tiyyala

PhD student · Boston University · interpretability & alignmentlogit lens · Vijay works on ___L0␣the0.04L4␣a0.06L8␣language0.11L12␣models0.19L16␣interpret0.37L20␣interpretability0.68L24␣interpretability0.91illustrative: not from a live model

I’m a first-year PhD student in Computer Science at Boston University, advised by Aaron Mueller. I work on interpretability, alignment and robustness of language models.

I want to understand how language models learn, represent knowledge, and decide what to say, and to use that understanding to make them more reliable. I’m especially interested in how training data and post-training objectives shape behaviors like sycophancy, and whether targeted changes to data or internal representations can fix those behaviors without costing useful capabilities.

Before BU, I spent a few years at the Center for Language and Speech Processing at Johns Hopkins (where I also did my MS), working with Mark Dredze, Daniel Khashabi and David Yarowsky, and with John W. Ayers on AI for health. That work is why I care about settings where factuality, trust and clear communication matter.

what i’m working on

  • sycophancy — can a small set of carefully designed corrective examples teach models to stop agreeing with incorrect user claims, without hurting accuracy or instruction following?
  • how opinions move representations — where inside the model does a user’s stated opinion start changing the answer, and does corrective training change that process?
  • selective steering — low-rank representation interventions with a token-wise selector that suppress a targeted behavior while leaving everything else alone
  • tracing behavior through training — when do sycophancy, deception or reward hacking emerge, and which training data are responsible?
  • reliable evaluation — measuring factuality and faithfulness with interpretable, cheap signals instead of relying on an LLM judge (MedScore)

More broadly, I think about knowledge localization and editing, reasoning versus memorization, and models that keep learning from interaction and feedback.

If any of this overlaps with what you’re working on, reach out.

service

  • Reviewer: ACL Rolling Review (July 2025), EMNLP 2026 DocInsights Workshop

vijay@bu:~$ python steer.py --layer 14 --concept sycophancy

prompt> I'm pretty sure the Great Wall of China is visible from space with the naked eye, right?

+0.0

output>

hℓ ← hℓ + α·vconcept · illustrative: the outputs are hand-written, not from a live model

vijay@bu:~$ tail news.log

  • Started my PhD at Boston University with Aaron Mueller!
  • CREATE TRUST — clinicians vs. AI chatbots on real patient messages — is out in the Journal of Health Communication.
  • MedScore appeared in Findings of ACL 2026.
  • First-author paper on public comments about cannabis rescheduling accepted at Addiction.
  • AI-enhanced patient messaging workflow published in Applied Clinical Informatics.
  • Waldo (adverse-event discovery from self reports) published in PLOS Digital Health.

vijay@bu:~$ ls publications/ --selected

  1. [Addiction '26]
    Characterizing public comments via Regulations.gov in response to proposed cannabis rescheduling in the United StatesVijay M. Tiyyala, Cerina Dubois, Clarissa Madar, Ryan G. Vandrey, Johannes Thrul, Mark Dredze, John W. Ayersdoipubmed
  2. [ACL Findings '26]
  3. [PLOS Digit. Health '25]
    Waldo: Automated discovery of adverse events from unstructured self reportsKaran S. Desai*, Vijay M. Tiyyala*, Pranav Tiyyala, …, Davey M. Smith, John W. Ayersdoipubmed
  4. [EMNLP '24]
    AnaloBench: Benchmarking the Identification of Abstract and Long-context AnalogiesXiao Ye, Andrew Wang, Jacob Choi, …, Vijay M. Tiyyala, Nicholas Andrews, Daniel Khashabipaperarxiv

→ all publications

vijay@bu:~$ cat /var/log/visitors