Want to read this magazine in your preferred language? You can translate the entire document using any modern AI tool (such as Gemini, ChatGPT, or Claude). Just follow these simple steps:
Hello everyone. Today, I want to share a journey that began more than thirty seven years agoāan ongoing adventure into a realm I call eyes-free information access. Simply put, eyes-free computing is the ability to access, interact with, and manipulate information using speech, particularly in hands-busy and eyes-busy environments. As computing becomes increasingly ubiquitous in our daily lives, transitioning from bulky desktops to smartphones to entirely invisible smart speakers, this concept has evolved from a niche accessibility requirement into a mainstream computing necessity.
To understand how far we have come, we must look back to the computing environment of 1989. That was the year I started my graduate studies at Cornell University. Most people casually assume the internet was born in the mid-1990s, but in 1989, it was already a living, breathing entity known as the NSF backbone. We had FTP, and HTTP was just being conceptualized and invented. Sitting in a computer lab back then felt like perching on a lightweight folding tray strapped to the top of a massive jet engine. The plane had not quite taken off yet, but the raw, unbridled power was humming beneath you, and you knew it was only a matter of time before it took to the skies. At that moment, I realized a fundamental truth that would guide the next three decades of my life: electronic information is inherently computable. It does not have to be permanently tied to a heavy machine on a desk; it can go everywhere with you. My driving motivation was to build systems that delivered information when you want it, where you want it, and exactly the way you want it.
As I developed these systems over the years, several key insights emerged. The first, and perhaps most foundational insight, is that electronic information is completely display-independent. When I write down my thoughts to create presentation slides, my internal ideas are transformed into visual representationsācolored marks on a background. On a physical piece of paper, those ideas are essentially dead until a human being actively looks at them and reconstructs the meaning in their own head. But when you move those ideas into the realm of electronic data, they become alive. They can be computed, searched, manipulated, and dynamically transformed into entirely different mediums.
Consider the simple, everyday example of a weather forecast. On a desktop monitor or a smartphone screen, you typically see a structured grid or a table of numbers: the days of the week aligned at the top, accompanied by high and low temperatures and little graphical sun or cloud icons. However, if you ask the Google Assistant, "What is the weather tomorrow?" it does not stubbornly read a table of numbers row by row. Instead, it synthesizes the data and tells you in a natural conversational tone, "It is going to be sunny and 75 degrees."
These are vastly different representationsāone highly visual and spatial, the other conversational and auditoryābut both are derived from the exact same underlying electronic form.
This realization led directly to my second major insight: how you speak things is fundamentally different from how you print or display them. Speech is not merely a secondary, bolted-on output; it is a first-class medium of interaction. This philosophy drove the creation of my PhD thesis, a system called ASTER (Audio System for Technical Readings), affectionately named after my first guide dog. ASTER was designed to provide the ability to read, browse, and interact with electronic documents auditorily, with a special focus on highly complex, technical math content. Mathematical formulas are notoriously difficult to convey through speech because they rely heavily on visual, spatial layouts like fractions, superscripts, and subscripts.
ASTER parsed these structures and used auditory cuesālike changes in pitch and pacingāto convey mathematical meaning structurally. I call it a niche application because not everyone wants to sit around listening to spoken calculus, but it was a magnificent project because it systematically exposed the deep underlying challenges of auditory user interfaces.
To make eyes-free access truly effective, you first need deep structural access to the information itself. Since a vast majority of the world's information at the time was locked in visual print mediums like PostScript and PDF, I moved to California and spent four years at Adobe. My work focused on extracting high-level document structures from PDFs, effectively freeing the text from its visual prison. Later, at IBM Research, I worked on the emerging vision of the multimodal web. We explored standards for multimodal computing, asking how we could represent information so it could be accessed seamlessly across different environments, whether the user was on a desktop, a Palm Pilot, or an early PDA.
Eventually, this brought me to Google. Initially, I worked on presenting search results in the form most convenient to the user. Then Android happened, opening up the world of mobile computing and suddenly turning the vision of ubiquitous, eyes-free access into a worldwide mainstream reality.
This brings us to the third critical insight: user interface peripherals ultimately determine the size, shape, and physical footprint of our computing devices.
What exactly is a user interface? At its core, it is simply the connection between a human and a machine. You, the human, express your intent by pushing a button, swiping a screen, or speaking a command. The machine computes this intent as data and presents a result in a way that grabs your attention. Because the core compute and storage functionalities can be entirely separated from the interface, the physical form of the device is dictated entirely by how we interact with it.
Think about the physical evolution we have witnessed over the decades. The desktop PC required a bulky monitor, keyboard, and mouse, permanently anchoring it to the corner of your office. Laptops freed us slightly, giving us portable LCD panels. Smartphones reduced computing to a shiny piece of glass that fits perfectly in our pockets. Today, smart speakers have made the computer virtually invisibleāit is just a microphone array waiting in the background. In the late 1990s, the IBM Microdrive arrived, and we marvelled at having one gigabyte of storage on a postage stamp.
We thought the future was carrying our data everywhere. Today, the data and processing live effortlessly in the cloud, yet we are still carrying our glowing displays. If you read Charles Dickens, you notice people in the 19th century constantly carrying candles and lanterns from room to room to see in the dark. Today, we carry our screens for the same reason. But as our networking becomes genuinely ubiquitous and our interfaces rely more on ambient voice and context, we will eventually stop carrying these displays. The next leap will likely involve invisible sensors integrated directly around our bodies.
The trajectory of this technology has profound and immediate implications for society. In the inclusive education think tank I co-founded with my colleague Ram Kamal, we have dedicated ourselves to ensuring that differently abled students can fully leverage this technological evolution. By breaking away from rigid, display-dependent paradigms, we can systematically eliminate traditional barriers to learning. We do not need to force students into singular, inflexible moulds. Instead, technology allows us to shape the environment and the information directly to the student, proving that with the right multimodal tools, differently abled individuals can achieve unprecedented independence, parity, and success.
As we look toward the future of eyes-free and multimodal computing, I am filled with immense optimism. At the end of my presentations and talks where ever, I often show a photograph of my guide dog sitting confidently in the pilot seat of a 767 aircraft. It serves as a humorous but poignant reminder: when we build smart, adaptable systems that conform to human needs, rather than forcing humans to adapt to the machines, the sky truly is the limit.