Sunday, May 27, 2012

Report: Apple CEO Cook performing well but Steve Jobs ‘would have lost his mind over Siri’

And you wind blows at Apple. $140 billion market value later after Cook took the helm at Apple, the company is going through some cultural changes. 

Steve Jobs would have “lost his mind over Siri,” according to an ex Apple employee.

The ex-employee told Fortune that people at Apple are “embarrassed by Siri”. Siri is Apple’s voice-recognition ‘PA’ feature on iPhones that has come under fire recently in a series of lawsuits that claim Siri doesn’t work as advertised. Siri was tagged as beta when Apple launched the iPhone 4S, and yet the company still picked it as the lead feature in its advertisements for the new phone.

Fortune has published a report that looks into the ways in which Tim Cook is changing Apple. Siri is noted as an example of a product that doesn’t reflect the normal quality of Apple products.

The report states: “The ultimate ‘tell’ of tectonic changes at Apple will be the quality of its products. Those looking for deficiencies have found them in Siri, a less-than-perfect product that Apple released with the rare beta label in late 2011, a signal that the service shouldn’t be viewed as fully baked.”

However, much of the report is very positive about Cook’s impact on Apple. It notes that both Apple employees and Wall Street analysts like and respect Cook, concluding that “as Apple enters a complex new phase of its corporate history, perhaps it doesn’t need a god as CEO but a mere mortal who understands how to get the job done.”

The report is full of examples of how Cook has changed the lives of his employees, who seem to have got back some of the work life balance that they didn’t have under Jobs. The tale of a meeting between a former Apple employee and a current Apple engineer is noted as one example. The former Apple worker assumed his friend would need to return to work following the meeting and was surprised that they had time for coffee. His conclusion: “I think people are breathing now.” However, Fortune notes, “it’s not necessarily a compliment.”

Another example of the more positive working environment is the case of a recent ‘Top 100′ company meeting – a tradition that gathers 100 of Apple’s top executives together for presentations – which was described as “upbeat and even fun,” in contrast to what it would have been under Steve Jobs. “Cook was said to be in a jovial, joke-cracking mood – a stark contrast to the grim and fearful tone Jobs engendered at the meetings,” notes the report. “Participants left the Top 100 energised about Apple’s near-term outlook.”

Fortune also reveals that Cook “often sits down randomly with employees in the cafeteria at lunchtime,” according to the Walter Isaacson biography, Jobs typically dined with Jonathan Ive.

It’s not just Apple staff that are seeing the friendly face of Cook. Cook is surprising investors by actually turning up to meetings and paying attention, something that Steve Jobs wouldn’t have done. According to the Fortune report, this is a “subtle but significant change” that means investors now have the CEO’s ear for the first time in years.

Cook is also a hit with Wall Street, not surprisingly since the company’s market value is up US$140 billion since Cook took over. “Wall Street in particular has good reasons – billions of them, actually – to love the Cook regime,” states the report.

Cook’s progress in China is also keeping analysts and shareholders happy. Fortune mentions Apple’s strengthened relations with China. In particular Cook’s visit to the Foxconn factory after Apple was criticised for working conditions there.

The report also notes Apple’s investments in China, suggesting that when Apple disclosed that it was US$2.6 billion assets there the number spoke to the value of the material and equipment Apple has bought on behalf of its suppliers. “Apple risks its own capital – with US$110 billion in cash at last count, it has plenty to risk – as a way of financing massive upgrades in its manufacturing capabilities in Asia, even though its partners will operate the equipment.”

As a result of Cook’s appointment as CEO, “Apple has become slightly more open and considerably more corporate,” according to Fortune.

One former engineering vice president told Fortune: “It looks like it has become a more conservative execution engine rather than a pushing-the-envelope engineering engine. I’ve been told that any meeting of significance is now always populated by project management and global-supply management. When I was there, engineering decided what we wanted, and it was the job of product management and supply management to go get it. It shows a shift in priority.”

Another change is with mergers and acquisitions. “Steve Jobs basically ran M&A for Apple,” notes the report. Today there is a department that allows Apple to “work on three deals simultaneously”. There are also more staff with MBAs being employed. Fortune notes that 2,153 Apple employees reference the term “MBA” in their LinkedIn profiles, and of those, “more than half the employees who reference “MBA” have been at Apple less than two years.”

The report suggests that “Cook is taking action that Apple sorely needed and employees badly wanted.” It also notes that “It’s almost as if he is working his way through a to-do list of long-overdue repairs the previous occupant (Jobs) refused to address for no reason other than obstinacy.”

Cook is “putting his stamp on Apple” and revealing “signs of Apple becoming a more normal company”.

“He seems, at the end of the day, to be honoring one of Jobs’ dying requests: That Apple’s management not ask “What would Steve do?” and instead do what’s best for Apple,” states the report, which can be read in full here.

via http://www.speechtechnologygroup.com/speech-blog - And you wind blows at Apple. $140 billion market value later after Cook took the helm at Apple, the company is going through some cultural changes.   Steve Jobs would have “lost his mind over Siri,” according to an ex Apple employee. The ex-employee told Fortune that people at Apple are “embarrassed ...

Friday, May 25, 2012

Is it possible to chat to a computer?

The holy grail of speech recognition is to have a completely conversational dialogue with a computer using speech recognition......

YOUR HORSE IS stuck in a ditch. You’ve tried luring it out with a sugar lump, but it’s not budging. Whipping out the smartphone, thumbs shaking, you Google “Annie”, the RSPCA’s virtual assistant. Reassured, you see Annie has handled this sort of thing before.

 “My horse has fallen down a ditch, what should I do?” is a Frequently Asked Question. A human operator might have offended by asking how your horse got there, whether you’d tried the sugar lump trick or some commentary from the Grand National? But Annie plays it safe: “Contact the RSPCA’s 24-hour cruelty and advice line immediately.” Why didn’t you think of that?

Chatbots such as Annie can only answer questions they’ve been scripted to answer. More sophisticated digital helplines can be part human, part automated, so you know your particular crisis has been understood and the advice given is appropriate. One day, though, fewer human ghosts will be needed in these machines.

David Levy is the organiser of this year’s Loebner Prize competition for the most human-like conversational computer, or chatbot. “The point is simply to create the ability in software to conduct an interesting, entertaining and informative human-like conversation which could be used in all sorts of different environments – commercial environments, social services environments, and just as entertainment,” says Loebner.

For the moment, this is something of a pipe dream – as I found out as one of this year’s Loebner Prize judges. It’s an annual competition based on a problem posed in 1950 by the great mathematician and second World War codebreaker Alan Turing. Many regard Turing as the father of Artificial Intelligence (AI). Can a machine think?

Turing suggested a test for a thinking machine. If, in a blinded, text-based conversation, a machine could convince a human judge that it too was human, then it could be considered a thinking machine. And so the famous Turing Test was born.

For 20 years, Hugh Loebner, a Hawaiian-shirted American philanthropist, has been footing the bill for a Turing Test-based competition. This year, to mark the centenary of Turing’s birth, it was held in the cloistered, wood-panelled confines of Bletchley Park, where Turing and his colleagues cracked the German Enigma code during the second World War.

Over two hours, we judges were to pit our wits against four different pairs of anonymous humans chatbots. Could we spot the difference?

You could almost hear the drum roll as we sat poised at our computers and, to a countdown Nasa would be proud of, launched into action. Judges initiated conversations. After that it was a free-for-all. We quickly realised the chatbots stuck out a mile and by Round three, I was getting silly …

Conversation A

Me: “Greetings earthling.”

Response: “Nanu nanu.”

Conversation B

Me: “Why did the chicken cross the road?”

Response: “I give up. Why?”

You might think that “nanu nanu” is an odd utterance for a human, but it had me laughing out loud, not something you’d expect from a chatbot. However, even a humourless human would surely know why the chicken crossed the road, wouldn’t they? B was clearly the chatbot. B was this year’s Loebner Prize winner, Chip Vivant.

Mohan Embar, Chip Vivant’s programmer, later tipped that if you’re ever in doubt about whether your online correspondent is human try asking this simple question: “where is the nose on your face? A human will think of the three-word answer (in the middle), but computers, even those with on-screen avatar faces, will never give this response.

Chatbots don’t actually understand us. They just pretend to, answering our questions based on sample human exchanges in their database. The bigger the database, the better they are at responding appropriately (some mine Twitter for this). The humanity of their answers and their ability to deflect or clarify difficult questions depends on the skill of their programmers.

Digital helplines are their most obvious application. But if they could be more human, would we prefer talking to them?

Surprisingly, not always. A 2007 survey by Callcentres.netfound that 67 per cent of Australians would rather deal with an Aussie-accented speech-recognition system than an offshore human call-centre agent.

If my experience judging this year’s Loebner Prize is anything to go by, humanity’s unique claim on thinking is safe for now. Like my three fellow judges (another journalist, an accountant and an archaeology professor), I concluded that a typing Orang Utan, like The Jungle Book’s King Louie could “be like me” more effectively than any of this year’s finalists.

One chatbot annoyed with a cacophony of questions. Another seemed rude and aggressive (a ruse used to stop difficult conversations mid-flow and revert to a more familiar subject). A third had the kind of verbal diarrhoea you dread from a dinner party guest.

None of the above is ideal in a helpline chatbot. But Chip Vivant, the winner, seemed to have potential and an important message from its maker – forget about apeing humans; be yourself; people will love you for it. Customers will love you for it. Given the current state of technology, Chip’s winning approach makes sound commercial sense.

Digital helplines must be fast and effective to satisfy customers. For now, says Erwin van Lun, founder and chief executive of chatbots.org(an online community of chatbot developers, academics and users), the human element is unnecessarily costly and risks doing more harm than good. “If your computer breaks, you want to know how to fix it. You don’t want to read something like ‘Your computer is broken? That’s a pity for you. Everything breaks sometime’.”

The most effective digital helpers are industry-specific and specialised, says van Lun. Few could pass the Turing Test, and whether or not they have a human face depends on the company’s branding strategy.

Anna, Ikea’s virtual assistant, has a face. So does Jon, Morrison’s straw-boatered virtual fishmonger. But Siri, the iPhone’s digital helper, doesn’t. And Quark, Creative Virtual’s “v-assistant” looks like an animated motorbike helmet.

And although in Ireland organisations seem slow to embrace automated helplines, they are bound to appear soon on a website near you.

A recent survey of the world’s top 1,500 companies by Gartner Research predicted that by 2015 half of all online self-search activities will be carried out by chatbots or “virtual assistants”. Company benefits lie in slashing 5 per cent off the cost of customer service, building brand loyalty and giving customers a bit of fun through playing with a robot.

Building charm into your chatbot can have customers coming back for more.

I found chatting to Cleverbot quite addictive. This is a common response, says its developer Roll Carpenter, twice winner of the Loebner Prize (2005 and 2006). Last year, Cleverbot won a Turing Test at the Techniche Festival in India, achieving a score of 59.3 per cent human.

Since then, Carpenter says, millions of people talk to Cleverbot every month for entertainment – the longest conversation on record lasting 11 hours.

“Despite it clearly being a bot,” people do regularly become convinced otherwise. Its because Cleverbot learns from people and imitates them … and has a huge variety of responses.”

Through his company Existor, Carpenter has two new projects on the go, a virtual assistant for a Russian bank and his cutest chatbot yet, Pupito, an animated puppy app for iOS. Pupito yaps and talks in text in response to typed input or, for a small fee, will chat with its “owner” in what Carpenter describes as puppy-like language (“Me is extra-specially clever, you knows!”).

But digital pets could have loftier aims than passing the time on the morning commute. They could be companions for people who, for some reason, can’t have the real thing. Elderly people in care homes, for instance.

Comforting the lonely is what Mohan Embar is all about. The US-born Indian programmer of this year’s Loebner Prize winner says that ELIZA is his inspiration. ELIZA was one of the worlds first ever chatbots, developed in the mid-1960s at MIT. It’s DOCTOR script simulated a form of psychotherapy which was so effective it completely fooled some of its “patients”.

“Imagine what we could do to provide comfort to people, using today’s technology,” says Embar. “That is my goal in life … Elderly people and those living in isolation, unable to obtain normal human interaction, could be helped by the “synthetic comfort” that an artificial agent could provide. I don’t know how rich this idea would make me, but the idea of being able to make an extension of myself that provides comfort to a lot of people is very appealing to me.”

And wouldn’t a bit of digital comfort be useful on a customer helpline, too?

via http://www.speechtechnologygroup.com/speech-blog - The holy grail of speech recognition is to have a completely conversational dialogue with a computer using speech recognition...... YOUR HORSE IS stuck in a ditch. You’ve tried luring it out with a sugar lump, but it’s not budging. Whipping out the smartphone, thumbs shaking, you Google “Annie”, the ...

New York Times columnist David Carr talks media

How a popular writer in 2012 puts technology to work, Not surprisingly speech recognition is a big part of it.......


New York Times media columnist David Carr
David Taintor May 25, 2012, 8:07 AM

David Carr, media columnist and culture reporter for The New York Times, talked recently to TPM about the future of news, the “glorious” view of the Port Authority Bus Terminal from his desk and what would happen if Twitter went dark tomorrow.

What is your writing/reporting day like?

Today, I spent far too much time on email and Twitter, and far too little time reporting a story I’m working on. I stayed up until 2 o’clock last night. I did sort of a personal elegiac post about Netflix, comparing it to my experience with AOL. I like to write at night. I’m kind of a vampire, and I have trouble writing in the newsroom. When I was first at the Times, I was pretty much on Business Day, cranking out stories, but I do better at home on the back porch or the kitchen table.

I wake up, four newspapers, usually get to two of them: New York Times and the Wall Street Journal. I also get the New York Post and the Star Ledger. I mostly look to Twitter to see what’s going on.

What does your media consumption look like on an average day?

The sort of hierarchy is: email is of the most interest to me; Twitter, second most interest; RSS, third most interest; and then I really don’t do any general web surfing. I don’t have a landing page. On my main Google page, there’s a bunch of widgets for like Ad Age and Poynter. My RSS is on there. Gawker is on there. I check The Guardian. Reuters is on my page. The Smoking Gun. A lot of Wired. Our own blog, Media Decoder. Salon is in there, that’s a recent addition. I think Salon is kind of coming along.

Can you answer the question you posed recently? Being a reporter is the ________ job in the world.

Well it’s the grandest caper there ever was. Really, if you can find work to be, sort of, professionally curious. I’ve always liked it, no matter how much I was working or not working. If you can find a way to make the economics work, it’s a pretty good way to go.

What’s the most significant journalism trend of 2012 so far?

This isn’t a trend yet. I have one of the new iPads. For both my email and Twitter, I’m able to talk into it and get speech to text (with Dragonfly voice-recognition software). It’s become more and more efficacious, which is great and convenient. If you start to think — one of the only barriers of people publishing is people have to type. What if that goes away? Just a huge explosion. Instead of going on the phone and talking about prom night, we’re either at or near a place where they can speak it in their phone and it’ll appear in text. What if people don’t have to type to get it into there? Does that make what we do more or less valuable? There’s more for us to sift through. But we’re the signal in the noise, and it’ll make us more valuable. I’m fairly democratic in my impulses. I don’t want more crap out there, but I think the fact that I need to type my thoughts means I share less of them.

Who’s doing good journalism on TV?

I used to work with Jake Tapper. I did a column that was off his reporting, about use of the Espionage Act. I just saw a very cool FRONTLINE thing on the financial crisis. I thought it was awesome. It made me think of the power of video and storytelling. I’m a Brian Williams watcher, but I’m usually still at work when he’s on. I’ve been setting the DVR on the weekends and watching Up with Chris Hayes. He’s sort of up to something else. I am working on a TV show, a web show, with A.O. Scott, and I made a little beta culture/media show. And it’s your basic, two doughy white guys talking into the camera.

As more and more TV sets become Web-enabled, I think the sort of boundaries — I’m in the radio business, you work at a blog, you work at a newspaper — whatever business you’re in, we all have text, video, little videos, big videos. It all sort of looks the same. I don’t think we’re in that different of a business.

What’s on your desk?

At work, I live on the fourth floor, which is in culture, and I’m in the corner. It’s a really glorious space. It looks out on the Port Authority (laughs), the Duane Reade. Right in front of me, I have a very large James Brown dancing figure and a lot of tin wind-up toys. There’s always coffee there, and there’s always Diet Coke and always a lot of crap.

At home, the coffee is always there and, more often than I’d like to admit, there’s an ashtray and a pack of cigarettes. I look out at a yard that could use some work and my bicycle.

Is there a general bad habit the news industry has right now?

The whole thing of, I put some top spin on something someone else did, and that constitutes work. That I put a little special sauce on top of somebody else’s get. It’s self-inviting, because that’s often what my column is: aggregation. But I like to think I phone call it in. I work it. I think it’s okay to aggregate and reframe within reason, just don’t pretend like it’s journalism.

The other thing is, now that there’s Twitter and the fights over scoops … no one cares. When the metrics of scoops get down to seconds and milliseconds, I don’t think it’s meaningful.

What do you think would happen to the news industry if Twitter went dark tomorrow?

I don’t think it would be a bad experiment. Twitter has a strong coastal bias. It’s got a strong media industry bias. When you’re in the Twitter dome, and you have something that blows up large, we grab onto it because it’s a place where we can see things pick up heat. But it’s sort of a false heat. It isn’t real. We’re all just shouting at each other. I got a look at my Twitter analytics, and I think if you start to guide yourself by what will trend within your group, you’re going to frequently end up in a tiny hole. But I love it. I’ve met real friends through it.

How will we get our news this time next year?

The rise of the visual Web is significant. I started working at Inside.com in 2000 and it was Michael Hirschorn who said, ‘That big picture and the small amount of text on your screen is your product. Everything else you write is gravy, it’s fine, but it’s not your product.” I do think people are going to be navigating in ever more visual ways. And I also think verbal navigation.

What do you tell recent journalism graduates or people trying to start a career in media?

I tell them that they should make stuff. The tools of production are at hand for everyone. I used to hire a lot of young people when I was the editor of Washington City Paper, and you used to have them show you the clips and see where else you worked. Show me what you’ve made with your own bare little hands. That, I think, is super important. People say, “You should’ve been here for the good old days.” I think that’s crazy. Yeah, it’s a little harder, but you have so many more tools at your disposal to story-tell. It’s cool to be in a business where you still learn. You don’t have to be able to code yourself, but you have to know what coding is. You should be able to work in Final Cut Pro. Wordpress should be second-nature. I think, in generational terms, being able to produce and consume content at the same time.

What’s the best advice you ever got?

A couple of things. One is, I think, Michael Ventura — I don’t know who he is or where he is — said: Writing is something that happens alone in a room. Regardless of the cacophony around you, and the mayhem, and the data stream, you have to remember to engage the piece of work you’re on fully, and remember that language is the biggest tool at your disposal. Why get in the business of telling stories unless you’re going to do a good job at it?

The advice I always gave young reporters is congruence. Don’t be super friendly on the telephone then hang up and just nuke somebody. If it’s a hard story, you should communicate that. Congruence between the way you report and the way you write is important.

What’s the most under-covered story right now?

Who is fighting distant wars. It’s being fought by the same people over and over. As war becomes increasingly remote and mechanized, the accountability that goes with that sort of goes away. The changing nature of warfare and what it does.

What’s the last book you read?

I’m right in the middle of Kurt Andersen’s new book, True Believers. He’s the one who brought me to New York at Inside.com. I’m excited to see how this one turns out. Next one I have up is Fathers Day, which is Buzz Bissinger’s new book.

Any plans to write another book? What would you write about?

I’m not a big threat to write a book. I found writing a book to be really hard, and I enjoyed it. But if you still have a job and you’re writing a book, it will really ruin your life. I’m really interested in video and television, and chatting on that. I’m sure that’s an effective way to ruin your life as well. But I’m naive about that, and maybe I could stumble in that. A nonfiction book, it would be like my book and my website smashed together.

© 2011 TPM Media LLC. All Rights Reserved.
via http://www.speechtechnologygroup.com/speech-blog - How a popular writer in 2012 puts technology to work, Not surprisingly speech recognition is a big part of it....... David Taintor May 25, 2012, 8:07 AM David Carr, media columnist and culture reporter for The New York Times, talked recently to TPM about the future of news, the “glorious” view of th ...

Harry Potter casts its spells on Kinect this fall

Ancient witch craft meets leading edge technology. Now you can use speech recognition technology to cast a spell.....

  Warner Brothers has announced that it will be making a new Kinect exclusive title ‘Harry Potter for Kinect’, planned to release this fall. You can check out a few screens and get the details below.
 
The announcement was made via press-release, the Xbox 360 exclusive was designed for the Kinect and features full HD graphics, motion-controls and voice-recognition. In a prepared statement sent to the press, Warner Brothers talks about the new title:

Based on all eight Harry Potter films, Harry Potter for Kinect allows players to join Harry Potter, Ron Weasley and Hermione Granger as they embark on an unforgettable journey through Hogwarts School of Witchcraft and Wizardry and beyond. With Kinect’s scanning technology, for the first time ever in a Harry Potter game, players will be able to scan in their own face to create a unique witch or wizard to journey through the adventures of the film.
 
In addition, Kinect’s hands-free and voice recognition capabilities allow gamers to cast spells using physical manoeuvres and by calling out spell names, offering an extraordinary, immersive Harry Potter experience.

“Harry Potter for Kinect will engage Harry Potter fans old and new by bringing them into the wizarding world as truly active participants,” said Samantha Ryan, Senior Vice President, Production and Development, Warner Bros. Interactive Entertainment. “Kids and parents will enjoy recreating their favourite Harry Potter adventures, from flying a broomstick in a Quidditch match, to battling pixies and duelling other wizards.”   The game will include some known playable characters as well as give players the chance to play as their own Avatar. The title will feature “some of the films’ most memorable moments”, though no specific story events were mentioned. The press-release did state that “visiting Ollivanders, choosing a house at Hogwarts, and confronting He Who Must Not Be Named in a climactic final battle” will be included in the title. The game will also feature a form of co-op play.
 

google.com by John Stewart
via http://www.speechtechnologygroup.com/speech-blog - Ancient witch craft meets leading edge technology. Now you can use speech recognition technology to cast a spell.....   Warner Brothers has announced that it will be making a new Kinect exclusive title ‘Harry Potter for Kinect’, planned to release this fall. You can check out a few screens and get t ...

From punch cards to speech, our input method methodologies - CWDN

The journey from punch cards to speech recognition as the input interface took us just a little more than 30 years.......

Where we once considered punch cards to be at the cutting edge of human computer interaction, the years passed and industry innovations brought us forward to a point where speech recognition technology has advanced to a point at which I am writing this blog without touching my keyboard.

Note: letters not forget the intermediary years in between punchcards and speech recognition, where we were perfectly happy with keyboards, mice and touchpads of various kinds.

After visiting Nuance in Boston last week and interviewing several of the company’s executives on the subject of speech recognition, it seems only fair to put Dragon NaturallySpeaking through its paces and speak this blog straight into a Word document to be later posted online.

Dragon is really pretty powerful now and although a few niggles will crop up in any spoken paragraph, I think that with a little training (of both the computer and myself the user) I could become far more used to using this input method - although I will have to get used to thinking as I speak rather than thinking as I type, which is not necessarily as easy as it sounds.

To extend my analysis of Nuance and its work with natural language understanding I also spoke to another vendor and so connected with Dr. Ahmed Bouzid, who is Angel’s senior director of product and strategy.

Note: Angel is a subsidiary of MicroStrategy and exists as provider of on-demand customer engagement solutions.

I asked Dr. Ahmed how the speech recognition software application developer community differs now compared to five years ago -
Is it easier to recruit now the talent pool is richer?

Dr. Ahmed Bouzid — Demand for speech scientists today far outstrips supply. The advent of Siri has been a galvanising event that has awakened the world to the possibilities of highly usable speech/voice user interfaces on the smart device. Evidence is the emergence of a whole new crop of Voice Assistants such as Evi, Cluzee, Eva, Ask Ziggy, and a couple of dozen more on all three of the three Mobile OS platforms (Apple, Android, and Micosoft). Available speech scientists (or software developers) today are indeed very difficult to find. I would say that the vast majority of them are taken up, not surprisingly, by Apple, Google, Microsoft, but also by AT&T and Nuance.

Who is wining in the speech vs touch battle?

Dr. Ahmed Bouzid — I think it is a mistake to view speech and touch as mutually competing interfaces. Speech is a highly compelling interface, but only in the right circumstances. You don’t want to use speech in a noisy place, or in a setting where you are not able to engage your device privately (e.g., financial transaction). On the other hand, when you are driving, you do not want to take your eyes away from the road — or your hands off your wheel. For that setting, speech is ideal. So, I would say that the winner is going to be whoever is able to understand that value is not inherent in any given interface but rather in how that interface is introduced in the user’s interaction stream. I would venture today that today, Siri is not there yet, nor are any of the Assistant Apps out there today. None of the Speech/Voice-enabled assistants today combine speech and touch in a compelling way that empowers the user to do what they want, they way they want it, and when they want it.

Will speech data now become part of the big data mountain in the cloud?

Dr. Ahmed Bouzid — Yes, indeed. The recorded voice is a highly rich collection of data points, and at least for now, such data is transferred over the network for processing. Since the arrival of Siri, Apple has collected billions of audio snippets from actual people asking actual questions (serious as well as silly). Google has done the same and in fact has developed a highly accurate speech recognition engine as a result of the audio data that it has collected over the last few years. Microsoft also runs its speech engine in the cloud and similarly has a treasure trove of audio data. Such data will not only enable these companies to continue refining their speech engines, but may also push these engines to be resilient enough to be highly accurate in data mining other audio (e.g., podcasts).

punch.bmp

Note: we still have some way to go with speech recognition, but the advancements that have been made make the technology really quite impressive and fascinating (I’m still talking not typing). We will still get problems with Homer names, homonyms — that’s better, when words do sound the same… As that live mistake just shows. But this has to be part of the way we start to use computing devices more in the future wouldn’t you agree?

via http://www.speechtechnologygroup.com/speech-blog - The journey from punch cards to speech recognition as the input interface took us just a little more than 30 years....... Where we once considered punch cards to be at the cutting edge of human computer interaction, the years passed and industry innovations brought us forward to a point where speech ...

Thursday, May 24, 2012

Knowing replaces searching - Educational Links

Speech recognition and search engine functionalities are growing closer together. Search is currently going through a major overhaul. After Google's latest penguin updates and Bing's changes we're looking at a complete different way of doing search.....

bing-logo

World wide webMany people consider the idea of computers communicating fluently with humans to be science fiction, but a large number of scientists and researchers see it as a soon-to-be reality instead. Looking back at the past century of technological innovation, it is obvious that society has come quite far in a short amount of time in regard to complexity of thought and the accessibility of information that students now enjoy. In the most recent decade, search engines like Google and Bing have been at the forefront of establishing and improving the human-computer information stream, allowing users from all over the world to share knowledge with one another in order to improve their educations. For instance, although search engines have been thus far confined to the random surfer model that calculates the probability of any one user clicking a link to a particular page, called PageRank, researchers are continuing to imagine and engineer new search strategies, and there is only a matter of time before the internet diverges from this aging search model. These search engines are not only being used in a surprising variety of ways but are now experimenting with new systems and techniques for data processing and retrieval that incorporate speech and thought in order to integrate man and machine for a more productive future.

Researchers and students in a variety of fields are especially well-poised to understand the potential impacts these new systems will have on the way society thinks and learns, many having benefited from these advances already. Instead of compiling long, detailed corpuses of words and expressions from physical documents as they have in the past, for instance, linguistic researchers are now running their data-analysis queries through platforms such as Google Books in order to more naturally determine such factors as common word and phrase usage as well as frequency of alternate spellings and grammar. Within the next five years, it would be no surprise to see most search platforms converting speech into search terms. To be sure, some platforms have already begun to do so, such as Apple’s Siri, which interprets speech as text for the iPhone’s search engine to use in information retrieval, then translating the most likely results from text to speech. Eventually, linguists can look forward to analyzing whole new corpuses of actual speech due to these processes, and everyone will be enjoying hands-free access to information as they go along with their daily lives.

bing-logoWhile Google has been a major innovator in the past when it comes to designing the best search engine, a new competitor has emerged in Microsoft’s newest search engine, called Bing. Instead of simply ranking the most popular websites in relation to users’ search terms like Google does, Bing attempts to read its users minds, mimicking such mental processes as free association and bridging the gaps between misidentifications and their targeted meanings. Because the platform is so new, however, Bing sometimes misfires or over-predicts users’ intentions. For example, when a user enters the search term “face,” Bing assumes that he or she is offering an incomplete search term associated with “Facebook,” and its first two links lead to the social network, while a Google search of the same term produces a results page that is more diverse. Obviously, the more search terms are included in a query, the more focused and accurate the results become for each of these engines, but they have yet to be able to truly predict human thoughts and intentions. Nonetheless, these developments are a step in the right direction—in a few more years, search engines will likely googlegrow more and more reliable in determining the intentions behind users’ search terms. In order to achieve this progress, innovators have already begun to identify a whole array of factors that may be used, including such intriguing concepts as tracking the past queries of a user and comparing them with search terms as they are entered as well as sci-fi-like face recognition software that can detect emotions and tailor results, and also advertisements, accordingly.

If many people in modern society take for granted the current level of the integration of information between humans and computers, the same people probably will continue to underestimate the significance of the multitude of advances to come to this field in the next several years. These advances will be indicative of the advance of civilization, freeing the human mind to reason at a higher level and to spend more time learning what to ask, instead of how to ask, concerning anything from a calculus problem to society’s most pressing questions. In the coming technological era, students will outperform their computers instead of the other way around, and internet search engines will provide society with the tools to do so.

Alex Petrovic – Brisbane, Australia. Dejan SEO Pty Ltd (more info) is a team of talented consultants and link builders. The company combines many years of experience in search engine optimisation and internet marketing among its consultants to provide clients with tangible results. Dejan’s 42-member SEO staff have skills and proven records in all areas of search engine optimisation including keyword targeting, competitor research, on-site optimisation, and link popularity. 

via http://www.speechtechnologygroup.com/speech-blog - Speech recognition and search engine functionalities are growing closer together. Search is currently going through a major overhaul. After Google's latest penguin updates and Bing's changes we're looking at a complete different way of doing search..... Many people consider the idea of computers com ...

Siri-like Dictation Features Coming to Mountain Lion? - TFTS

Windows users have access to speech recognition for many release versions. But it never seemed to have really taken off on the Windows desktop. 

Will the Mac now get a Siri version included with the new mountain lion OS, and will  this revolutionize speech recognition on the desktop?

Safari File Suggests Siri Is Coming to Macs in the Near Future, Sort Of
Posted on TFTS by on Wednesday, 23. May 2012 under: Computers & Computing, Mobile/Cell Phones, PCs, Notebooks, Netbooks & Tablets

Apple’s Siri virtual assistant is going to play a major part in current and especially future iOS devices, but the company appears to also be interested to include certain Siri features in its future OS X versions.


Siri is currently an iPhone 4S exclusive feature, although other iOS devices can unofficially run Siri. In a surprising move, Apple did not offer Siri support to new iPad buyers, although a Dictation feature is present on the 2012 iPad, which seems to suggest that Siri could soon be available on the tablet.

A recent report revealed that Siri will be enabled on the current iPad once iOS 6 is launched and claimed that developers will have access to Siri APIs in the near future. Siri is also rumored to be preloaded on the company’s first television sets, although these devices are yet to be confirmed.

More interestingly, a full-fledged Siri client could be available on OS X devices at some point in the future. Meanwhile, we’ll just have to make do with the Dictation feature – currently available on the iPhone 4S and iPad 3 – that’s reportedly coming to OS X Mountain Lion computers. According to 9to5Mac, Siri-like Dictation features have been discovered in OS X Mountain Lion:

According to a resources file inside of the latest build of Safari in the newest seed of the upcoming OS X Mountain Lion, Dictation might be making its way to Macs next. Since Macs do not sport virtual keyboards or physical keyboards with a microphone-labled key, users (by default) will apparently need to simultaneously click both command keys to start voice input.

However, other Dictation references have not been found in current Mountain Lion beta versions and the feature itself is non functional at this point. OS X Mountain Lion and iOS 6 are going to be the main stars of WWDC 2012, maybe right alongside new MacBook Pros, so in case voice-based features are coming to Macs we may hear more details about them straight from Apple in a few weeks from now.

Credit: Source. 

via http://www.speechtechnologygroup.com/speech-blog - Windows users have access to speech recognition for many release versions. But it never seemed to have really taken off on the Windows desktop.  Will the Mac now get a Siri version included with the new mountain lion OS, and will  this revolutionize speech recognition on the desktop? Safari File Sug ...