Andrew Shin
2026
Large Language Models’ Internal Perception of Symbolic Music
Andrew Shin | Kunitake Kaneko
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Andrew Shin | Kunitake Kaneko
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Large language models (LLMs) excel at modeling relationships between strings in natural language and have shown promise in extending to other symbolic domains like coding or mathematics. However, the extent to which they implicitly model symbolic music remains underexplored. This paper investigates how LLMs represent musical concepts by generating symbolic music data from textual prompts describing combinations of genres and styles, and evaluating their utility through recognition and generation tasks. We produce a dataset of LLM-generated MIDI files without relying on explicit musical training. We then train neural networks entirely on this LLM-generated MIDI dataset and perform genre and style classification as well as melody completion, benchmarking their performance against established models. Our results demonstrate that LLMs can infer rudimentary musical structures and temporal relationships from text, highlighting both their potential to implicitly encode musical patterns and their limitations due to a lack of explicit musical context, shedding light on their generative capabilities for symbolic music.
A Zipfian Analysis of Visual Token Distributions for AI-Generated Images
Andrew Shin
Proceedings of the 4th Workshop on Advances in Language and Vision Research (ALVR)
Andrew Shin
Proceedings of the 4th Workshop on Advances in Language and Vision Research (ALVR)
The rapid evolution of text-to-image generation has blurred the perceptual boundary between natural and synthetic imagery. However, it remains questionable whether the statistical structure of generated visual content mirrors the information density of the physical visual world. Drawing upon principles from statistical linguistics, this study investigates the visual language of generative models through the lens of Zipfian dynamics. By analyzing a large-scale corpus of real and synthetic images, we uncover a fundamental divergence between visual syntax and semantics. We find that while generative models have successfully replicated the low-level physics of light, their high-level texture vocabulary exhibits distinct statistical signatures. Our analysis reveals a spectrum of entropy, identifying architectural fingerprints unique to each model. Furthermore, we investigate the relation ship between generated images and prompt complexity, and find that increasing the semantic specificity of text prompts systematically degrades the statistical realism of the generated output.
2021
Transformer-Exclusive Cross-Modal Representation for Vision and Language
Andrew Shin | Takuya Narihira
Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021
Andrew Shin | Takuya Narihira
Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021
2015
Context-Dependent Automatic Response Generation Using Statistical Machine Translation Techniques
Andrew Shin | Ryohei Sasano | Hiroya Takamura | Manabu Okumura
Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
Andrew Shin | Ryohei Sasano | Hiroya Takamura | Manabu Okumura
Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies