Perceptual learning is a key component of human development. In the speech domain, for example, infants’ perception becomes specialized for the vowels and consonants of their native language(s) even before they can speak, a process known as early phonetic learning1,2. Early perceptual learning occurs in multiple domains—including speech, sign gestures, music, face, shape and motion perception2–7—and has far-reaching implications—including permanent perceptual deficits for infants with atypical exposure6–9, with cascading consequences for downstream cognitive processes10,11. Despite its pervasiveness and impact, the mechanisms underlying early perceptual learning are still poorly understood. Even basic questions remain largely unanswered. For example: Is learning generative or contrastive? Reconstructive or predictive? What is the nature of learned representations - clusters, continuous dimensions or more structured representations? Does learning involve encoding a form uncertainty or even full probability distributions? Does it require deep, iterated transformations, or are shallow maps sufficient? Does it involve explicit adaptation processes? What is the role of rehearsal versus online learning?
Testing fine-grained hypotheses about learning mechanisms through controlled learning experiments in the lab has proven challenging. Indeed, stringent operational and ethical constraints on infant experimentation severely limit the scale and scope of learning in these experiments12–14. Studying infant perceptual development at home as a form of large-scale natural learning experiment provides a promising alternative. For example, in the case of early phonetic learning, psychologists have documented dozens of language-specific developmental trajectories over the past forty years2,15. These could inform and constrain the set of learning mechanisms possibly at play. However, to do so, requires a reliable method to link hypotheses about learning mechanisms to their observable consequences on the timescale and under the conditions of natural perceptual development. This has been a key bottleneck.
In this PhD, you will leverage progress in machine learning/AI technology to overcome this bottleneck. Specifically, you will link learning mechanisms and observable predictions in early phonetic learning by explicitly simulating the learning process over developmentally-relevant timescales and with realistic input. Building on promising initial results16,17, you will refine this approach into a rigorous scientific methodology. A core conceptual challenge will be to organize the space of self-supervised representation learning algorithms so as to pick a finite set of learning simulations that best support (i) conclusions about the constraints placed on learning mechanisms by the available empirical record, and (ii) the identification of novel decisive experiments - where the available record might not suffice. A core practical challenge will be to develop, automate and scale processes to curate training and test data in multiple languages, and to train and evaluate models for multiple representation learning algorithms.
Depending on your interests, there are many possible extensions and openings are possible, for example to other language phenomena beyond phonetic learning (e.g., prosodic learning or word learning), to other modalities beyond sound and speech (e.g., face or other animated object perception in videos) or to other types of empirical evidence beyond behavior (e.g., EEG/MEG or fMRI recordings while participants listen to natural speech).
Requirements
The ideal candidate has a quantitative background and prior experience with representation learning algorithms. They have an interest for cognitive (neuro)science and are eager to rigorously develop innovative scientific methodology in the challenging field of computational cognitive science.