MQUD
MQUD contains 1,250 questions evoked by scientific figures, including 708 annotated by the papers’ original authors. We use it to study whether vision-language models can ask questions about a figure’s role in a paper.
Yating Wu & collaborators
MQUD contains 1,250 questions evoked by scientific figures, including 708 annotated by the papers’ original authors. We use it to study whether vision-language models can ask questions about a figure’s role in a paper.
ContextWeaver links an agent’s reasoning steps to the earlier steps they depend on. It uses these dependencies to select and summarize context for subsequent actions.
QSalience predicts how much answering a reader’s question would help them understand a text. It is trained on 1,766 context–question pairs annotated for salience.
QUDeval assesses generated discourse questions against four linguistic criteria, using human annotations to evaluate both parsers and automatic metrics.
ElabQUD pairs about 1,300 elaborations in simplified text with the implicit questions they answer. Modeling these questions improves elaboration generation.
We train a parser on DCQA to connect sentences through the implicit questions they answer. The resulting QUD dependency structures describe how a document develops its content.
QUDsim measures similarities in discourse structure between documents. We find that LLMs reuse these structures across texts, even when the topics differ.
We use language model surprisal and linguistic features to predict disfluencies in monolingual and bilingual speech, and compare the patterns associated with each group.