ARTICLE · TODD KELSEY
Helping Students to Become AI Bilingual - with Personal Data

There is increasing evidence that there is an interdisciplinary need in higher education to help students learn how to work with data in their disciplines. The President of MIT, Rafael Reif, called on the United States this past Monday to take AI seriously, in government and education, and rightly urges schools and colleges to help students to become “AI bilingual”, which could also be referred to as developing “AI fluency”. Since the current predominant form of AI, machine learning, relies heavily on vast amounts of data, developing fluency in data is a related issue, and it is critical to develop ways to help students from all background to get more comfortable working with data, in whatever discipline they are.
This is what MIT is seeking to do in their newly announced interdisciplinary AI College - and the evidence suggests that as many schools as possible should explore ways to help students get experience in working with data, and ultimately consider how data science and artificial intelligence may provide new opportunities in their careers - and also to develop a healthy respect for the economic disruption and impact AI and automation is having, and will continue to have, on jobs. The fundamental reality is that robots can be physical, and assist or replace humans; and the same is true for software robots. In both cases, automation is replacing and assisting many jobs, and creating new ones.
The question of how to help educate the nation about AI is truly an interdisciplinary one.
For the computer science student who has a background in statistics, probability, linear algebra, and programming, the transition to working with machine learning is relatively seamless - and though it may take time, there are certainly ways to help anyone learn more about these foundational tools. But in the classic leapfrog game of technology and business, tools and platforms are arising that attempt to make it easier for others to engage in working with data, without necessarily having to know how to create new algorithms for example. Companies such as Cognitive Scale and the Cortex tool represent this phenomena - giving visual designers the opportunity to work on designing neural networks.
Other platforms such as DataRobot, Google’s AutoML, and Amazon’s Sagemaker, can help companies to leverage machine learning without necessarily having an advanced data scientist on board - in fact, AutoML was borne out of necessity; as the severe shortage of AI talent continues; AutoML seems to let “AI develop AI” or “machine learning develop machine learning”. (And the related implication is - if something as complex as machine learning can be automated - is there anything that can’t be automated?)
And with the rise of assistive tools and platforms, the debate continues about whether it is better to know AI as well as a computer scientist or if you can truly make progress with a so-called “black box” - a platform that simplifies things for the user but that is not customizable. The debate would be a distant cousin to the philosophy of website designers since 1995 or so, where some believe you should be able to code directly in HTML, and others believe that using one of the assistive platforms like Weebly is fine.
Machine learning, AI and big data is arguably much more complex than developing a web page, but the same dynamic exists, and the answer is that both sides are right. Yes it would be ideal if everyone working with machine learning had the knowledge of a PhD in computer science, but significant progress and value can be added by people from other backgrounds, especially as education and assistive platforms catch up to the immense demand.
And meanwhile, what about all the different disciplines? I think each discipline has something to contribute. As Walter Isaacson concludes in The Innovators, the most successful innovation ecosystems involve a significant amount of cross pollenation from various functional areas and disciplines.
Case in point, even PhDs in data science have sometimes found it difficult to hit the ground running when hired in business, exactly because they have been in a silo of computer science, and may or may not understand how business works; the different dynamics, or just the way that data flows in a company. Certainly the disciplines of business and computer science have a lot to learn from each other; but it doesn’t stop there.
And how do you cast the net as wide as possible, and be inclusive, and help students and teachers and professors for that matter to “get comfortable” with data. How can you motivate and inspire someone who says: “I am not a person”. Or a student who simply wants to know where to start?
One way would be to start with life data: that is, helping a student from any discipline to start their journey in developing data fluency, through considering their own data: to recognize things like pictures and text constitute data. Conversations could begin with social media, and the issues surrounding privacy: the cultural, ethical, philosophical issues, as well as the practical reason why companies are so very interested in their data.
But closer to home, and perhaps with deeper significance, it may be that students’ own life stories, the heritage of their family and communities, could be an excellent starting point. It is life data - it might even include digital artifacts and data from social media, but it is a story. Considering your own life story, interviewing a family member, considering your heritage, including ethnic and cultural, can be deeply significant, even if it doesn’t come naturally at first. It is still true and highlighted to me by the recent wildfires in California, that one of the first things people might grab when they flee a burning house could be a photo album.
AI expert Pedro Domingos guides readers from any background in a masterful journey of learning about AI, in his book The Master Algorithm. That single book alone might be enough to kindle the reader’s imagination, as he openly invites readers or every background to learn more about AI and to go from being an observer to becoming a participant. Because the stakes are high and the opportunity is considerable, to leverage AI in every discipline, from health to science and business and every other field, to help make the world better.
And the author concludes that one of the biggest opportunities in AI is working with life data - helping people recognize what it is, to take control of it, and to manage it, in the emerging ecosystem of AI.
So my belief is that life stories and life data may be a good starting point; when you consider the need for finding ways to help students in a variety of disciplines to learn more about data, and when you also need to find ways of teachers and professors in different departments to work towards helping students develop fluency in data and AI.
I am in the midst of forming my own viewpoint about how to go about it, and I recognize that not every school, teacher or student for that matter is convinced that AI is that critical of an issue, even China declaring itself an AI-first country, and Google declaring itself an AI-first company, the President of the United States authorizing massive funding for AI (this past monday), or the fact that MIT is developing the first AI College.
Accordingly, part of the need is to start the discussion. I wrote the book Surfing the Tsunami for this very reason, and I am making it free to any educator, school, and all their students, or any legislator, in any country. It attempts to introduce AI, includes a fair mount of data (to try and convince readers to take AI seriously). And regardless of what you believe about how much disruption will occur from AI, it is clear that the issue is gaining prominence, and that it will continue to be an important issue to consider. For anyone interested please visit http://tsunami.ai
(Part of the challenge of developing the book is that things change so fast, but I am planning on updating it every year; and I believe it could be a useful resource in your own journey, or the journey of your school or company or government for that matter. I recognize optimists, and alarmists, and advocate for a realist perspective.)
Invitation #1: I think more books and curriculum are certainly needed, including guided tours through some of the many existing resources that can help people learn more about data, programming and artificial intelligence. If anyone is interested in telling me about a resource; or is interested in developing one, please let me know.
Invitation #2: And if you are interested in thinking more about the question or “life data” and “life stories”; please visit these two websites and get in touch.
The first is called http://glia.net - it is an interesting new project from a Xoogler (ex Googler) Richard Whitt. He is doing some good thinking about the nature of data, the need to help people take control of their data.
I also invite you to take a look at a past research project called Personal Digital Archaeology, which I am dusting off and looking at with new eyes, as a potential basis for helping students and communities to think about their stories, and their heritage, as a first step towards learning about data, and beginning to come to grips with the importance of personal data, corporate data, and the relationship between the two, as well as data in the sphere of government, science and every other field. http://digitalarchaeology.org
The spirit of the Digital Archaeology project is to recognize that life stories on digital media are in danger of disappearing, since digital formats evolve and change so often. And the idea is simply to encourage people to consider their life stories, their heritage, and how to capture, preserve and share.
For example, this picture is precious to me, a picture of my Grandpa Miller, normally a very distinguished man, who I convinced to put on a tie-dye one time and put my guitar over his shoulder.
And Grandpa was responsible for inspiring me originally to think about digital archaeology. When he passed away, I inherited a computer of his, with disks that contained his writing, which I couldn't access because they were obsolete. And I realized precious digital artifacts are slipping into oblivion, and that it would be a worthy cause to help people rescue, preserve and share them.
And as of today I am convinced that this kind of exploration could be a nice way to start exploring life stories, life data, and data in general, for learners of any age or background.
If you’re interested in considering how life stories or life data might form the basis for the goal of developing fluency in data and ultimately becoming “AI bilingual”, please get in touch.