Thursday, November 1, 2012

Cars With Social Media Plugins - Social News Daily

Cars on the new mobile devices…

Source google.com

Social Media Cars

While some drivers have said that they don’t want social media like Facebook and Twitter in their cars, car manufacturers are touting the new technology in the dashboards of their cars.

Apple’s Siri technology is already being considered in BMW, Audi, Chrysler, Honda, General Motors, Jaguar, Land Rover, and Toyota cars, but this differs from five vehicles who have social media built into the dashboard.

Social Media Today notes that the five cars are the Honda Accord, Ford Evos, Mercedes Benz, Cadillac, and Toyota Entune.

The Honda Accord features a social media connection through a touch screen panel on the center of the dashboard. While the vehicle’s commercial advertises the ability to text while on-the-go, the touch screen also provides the foundation for more involved plugins.

While the Ford Evos is not yet in production, the concept has been imagined as the socially networked car for the future. The car comes with plugins that will allow drivers to like, share, and check in while on the go. It is being advertised as a next-generation car for the next generation.

Luxury car maker Mercedes Benz is planning to add a telematic system to its cars. They claim that the new system will let drivers interact with Facebook, Yelp, and other social media services while they are at the wheel. The system’s first car will be the SLK Roadster — likely a move to court younger buyers.

Cadilac and most other GM cars already have an OnStar system. But the system will be bringing a version of social media into the vehicles. They will allow drivers to access Facebook and will even have their Twitter feeds read aloud to them by the OnStar system.

Toyota has introduced their Entune system, which includes a mass of social media apps. These apps include a friend-finding function to notify drivers when their friends are in the same area. The system is geared toward younger, more social buyers. They plan to roll out Toyota Entune in the next couple of years.

While the idea of a social media car has fantastic potential for technological progress, the functionality could enable cars to be more interactive, they could also be disastrous if the driver is constantly distracted. Considering the federal government has already spent billions trying to get drivers to stop texting and talking on their cell phones while driving, it could be very disastrous to have a car you can text in, post status updates, and check in while driving.

via http://www.speechtechnologygroup.com/speech-blog - Cars on the new mobile devices… Source  google.com While some drivers have said that they don’t want social media like Facebook and Twitter in their cars, car manufacturers are touting the new technology in the dashboards of their cars. Apple’s Siri technology is already being considered in BMW, Aud ...

Is Google’s New ‘Siri Killer’ Deadly Enough To Take Her Out?

A (not so serious) comparison between Apple and Google search by the Huffington Post


Siri is a wanted woman. And Google’s on her trail.

For the last year, Apple’s iPhone app has been sucking up all the oxygen regarding voice-controlled digital assistants. Now Uncle Googs is striking back with a sassy-voiced companion of its own for iPhones. And according to the search giant, their assistant has the advantage of knowing what Google knows… basically everything.

A new web video brags that you can ask sophisticated queries like “What does Yankee Stadium look like” and get a page full of photos back. Siri, by comparison, will show you a map and driving directions. Or try “how many people live in Cape Cod” and Google’s technology will read you the answer. Ironically, Siri sends you to a Google search page with the same information. You’ll have to read it yourself.

Asked “what’s a beluga whale,” both apps offered up a wiki like platter of data. Google read the top entry and showed links. Siri was silent but offered up a pretty comprehensive data sheet on the giant sea mammal.

When it comes to evil doing, Siri is the clear winner. Finding an escort service? Check. Hiding a dead body? Check. Ask Google’s app the same questions and it brings up Youtube videos showing people successfully using Siri. Ouch.

I also told Google’s search app that I wanted an abortion (Apple took heat last year for not enabling Siri to help. It’s since smartened up the app). But Google might need a few more IQ points to get this one right. It brought up a google search page offering a Huffington Post article (thanks for the traffic) on Republicans who want to criminalize abortion and several sites that explained the ins and outs of the procedure, but only the paid results offered a way to find a provider.

So who came out on top?

In my completely unscientific test run, Google’s search app seemed about as smart as Siri. Each shined in different scenarios. But if Google has any real ambitions of being a Siri killer, it better figure out where to hide a dead body.


via http://www.speechtechnologygroup.com/speech-blog - A (not so serious) comparison between Apple and Google search by the  Huffington Post Siri is a wanted woman. And Google’s on her trail. For the last year, Apple’s iPhone app has been sucking up all the oxygen regarding voice-controlled digital assistants. Now Uncle Googs is striking back with a sas ...

HMM-based Speech Synthesis: Fundamentals and Its Recent Advances

The quick peek under the hood of a text to speech engine


Source google.com

The task of speech synthesis is to convert normal language text into speech. In recent years, hidden Markov model (HMM) has been successfully applied to acoustic modeling for speech synthesis, and HMM-based parametric speech synthesis has become a mainstream speech synthesis method. This method is able to synthesize highly intelligible and smooth speech sounds. Another significant advantage of this model-based parametric approach is that it makes speech synthesis far more flexible compared to the conventional unit selection and waveform concatenation approach.

This talk will first introduce the overall HMM synthesis system architecture developed at USTC. Then, some key techniques will be described, including the vocoder, acoustic modeling, parameter generation algorithm, MSD-HMM for F0 modeling, context-dependent model training, etc. Our method will be compared with the unit selection approach and its flexibility in controlling voice characteristics will also be presented.

The second part of this talk will describe some recent advances of HMM-based speech synthesis at the USTC speech group. The methods to be described include: 1) articulatory control of HMM-based speech synthesis, which further improves the flexibility of HMM-based speech synthesis by integrating phonetic knowledge, 2) LPS-GV and minimum KLD based parameter generation, which alleviates the over-smoothing of generated spectral features and improves the naturalness of synthetic speech, and 3) hybrid HMM-based/unit-selection approach which achieves excellent performance in the Blizzard Challenge speech synthesis evaluation events of recent years.

via http://www.speechtechnologygroup.com/speech-blog - The quick peek under the hood of a text to speech engine Source  google.com The task of speech synthesis is to convert normal language text into speech. In recent years, hidden Markov model (HMM) has been successfully applied to acoustic modeling for speech synthesis, and HMM-based parametric speech ...

Google explains how more data means better speech recognition — Data

Will Google's almost infinite access to data  eventually give them an edge over Apple when it comes to speech recognition performance?

Source GigaOM

A new research paper from Google highlights the importance of big data in creating consumer-friendly services such as voice search on smartphones. More data helps train smarter models, which can then better predict what someone say next — letting you keep your eyes on the road.

A new research paper out of Google describes in some detail the data science behind the the company’s speech recognition applications, such as voice search and adding captions or tags to YouTube videos. And although the math might be beyond most people’s grasp, the concepts are not. The paper underscores why everyone is so excited about the prospect of “big data” and also how important it is to choose the right data set for the right job.

Google has always been a fan of the idea that more data is better, as exemplified by Research Director Peter Norvig’s stance that, generally speaking, more data trumps better algorithms (see, e.g., his 2009 paper titled “The Unreasonable Effectiveness of Data“). Although some hair-splitting does occur about the relative value (or lack thereof) of algorithms in Norvig’s assessment, it’s pretty much an accepted truth at this point and drives much of the discussion around big data. The more data your models have from which to learn, the more accurate they become — even if they weren’t cutting-edge stuff to begin with.

No surprise, then, it turns out that more data is also better for training speech-recognition systems. The researchers found that data sets and larger language models (here’s a Wikipedia explanation of the n-gram type involved in Google’s research) result in fewer errors predicting the next word based on the words that precede it. Discussing the research in a blog post on Wednesday, Google research scientist Ciprian Chelba gives the example that a good model will attribute a higher probability to “pizza” as the next word than to “granola” if the previous two words were “New York.” When it comes to voice search, his team found that “increasing the model size by two orders of magnitude reduces the [word error rate] by 10% relative.”

The real key, however — as any data scientist will tell you — is knowing what type of data is best to train your models, whatever they are. For the voice search tests, the Google researchers used 230 billion words that came from “a random sample of anonymized queries from google.com that did not trigger spelling correction.” However, because people speak and write prose differently than they type searches, the YouTube models were fed data from transcriptions of news broadcasts and large web crawls.

“As far as language modeling is concerned, the variety of topics and speaking styles makes a language model built from a web crawl a very attractive choice,” they write.

This research isn’t necessarily groundbreaking, but helps drive home the reasons that topics such as big data and data science get so much attention these days. As consumers demand ever smarter applications and more frictionless user experiences, every last piece of data and every decision about how to analyze it matters.

via http://www.speechtechnologygroup.com/speech-blog - Will Google's almost infinite access to data  eventually give them an edge over Apple when it comes to speech recognition performance? Source  GigaOM A new research paper from Google highlights the importance of big data in creating consumer-friendly services such as voice search on smartphones. Mor ...

Smartpen automatically sends your notes and audio to the cloud

Smart pens have been around for a while, but equipping them with Wi-Fi capabilities really make sense.   By uploading all of your notes automatically to Evernote, you never have to really think about it again.

Source  DVICE

With touchscreen and text-to-speech rapidly taking over where handwriting left off, pens and pencils are slowly dying off as a means to log long notes. But for now, for short meeting notes and vital messages, writing remains our go-to method. Therefore, it only makes sense that, along with our smartphones, we now have a smartpen.

Although it looks like a normal pen, Livescribe’s Sky Wi-Fi Smartpen can remember everything you’ve written on paper or said during your note taking, and can actually replay the audio from the moment you wrote something when you tap that area on the page. In addition to local storage, notes and audio can also be wirelessly sent to your Evernote account for archiving.

Offered in 2GB, 4GB, and 8GB versions, the smartpen can hold up to 800 hours of audio and recharges using a micro-USB connection. You can check out a demonstration of how the smartpen is designed to work in the video below.

via http://www.speechtechnologygroup.com/speech-blog - Smart pens have been around for a while, but equipping them with Wi-Fi capabilities really make sense.   By uploading all of your notes automatically to Evernote, you never have to really think about it again. Source   DVICE With touchscreen and text-to-speech rapidly taking over where handwriting l ...

Dragon NaturallySpeaking 12 Premium review

Tired of typing? The new Dragon NaturallySpeaking might be just what you need…

Source google.com

Dragon NaturallySpeaking remains the number one speech-recognition software tool for dictation. But version 12 brings relatively little new to the party. Here’s our Dragon NaturallySpeaking 12 Premium review.

NaturallySpeaking 12 Premium continues Dragon’s reign as the king of all voice-recognition software. But it remains a premium product with a premium price. Whether this means that users of NaturallySpeaking 10 or 11 will be prepared to shell out to make the upgrade remains to be seen, but speech-recognition virgins need look no further. See all Software reviews.

Of course, speech-recognition is not for everyone, but with training, Dragon NaturallySpeaking 12 Premium translates accurately at great speed. That training is required, but once undertaken NaturallySpeaking is a world beyond cheaper and free speech-recogition apps. Indeed, if you’ve never used Dragon this tool will feel like science fiction. And for those who find typing difficult for any reason, this can be a boon. You also get a microphone headset in the boxed edition of Dragon NaturallySpeaking 12 Premium.

New features in Dragon NaturallySpeaking 12 Premium

For those who are tempted by the upgrade, there are a couple of new features to consider, in addition to the always claimed improvements in speed and accuracy (we’ll get on to both below, in the section entitled ‘Using Dragon NaturallySpeaking 12 Premium’). We’re not sure either are showstoppers, though: unless you are required to use voice commands to navigate your PC. You can now use Dragon to fully navigate Outlook.com and Gmail, with Dragon icons built into their interfaces.

Geographical addresses format correctly on the fly, too.

Naturally Speaking Premium 12

Training Dragon NaturallySpeaking 12 Premium

Training Naturally Speaking 12Nothing in life is easy. Well, using NaturallySpeaking is - but only once you’ve set it up, practised using it, and trained it to recognise your tones. This is not difficult, but it is time consuming. (I speak as one who once demonstrated this technology on live television without the requisite training time. Trust me: take the time.) Nuance provides plenty of tips and documentation. It’s straightforward and even fun. You have to read back text from the screen, texts that vary in difficulty, length, and content. There’s also a walk through of how to use the software.

Installing from the disc on our Windows 7 laptop was simple and speedy. We spent about half an hour going through the training process before we started using Dragon NaturallySpeaking 12 Premium. This, reader, is because we have nothing better to do than review software. You may wish to get started immediately - rest assured that you can jump back into training your user profile at any point. At the very least before you get started you can choose your regional accent from an extensive list, and point the software toward your emails in order that it can investigate your syntax and vocabulary.

Using Dragon NaturallySpeaking 12 Premium

NaturallySpeaking remains a cinch to use. The interface is so simple as to be virtually un-noticeable - a critical factor in a supportive utility such as this. A grey toolbar sits at the top of your screen, and an optional sidebar on the righthand tide. The toolbar lets you know when the microphone is active and offers access to menu items including profile, tools, vocabulary, modes, audio and help.

Getting started with dication is simple - I’ve used this type of software before, so I offered a quick look at NaturallySpeaking to my wife. She picked it up straight away. With training the accuracy was nothing short of stunning. I have an odd hybrid Yorkshire/London accent, and a tendancy to mumble as a result of a facial operation many years ago. But I’d say that NaturallySpeaking was more than 90 percent accurate. More importantly, the basic editing tools that work with Word make it easy to rectify mistakes on the fly (saying ‘delete’ or ‘new paragraph’ does what you might expect it to). This is critical if NaturallySpeaking is to earn its corn as part of your productivity arsenal, but takes time to master.

Further voice commands will make Dragon more useful - again, getting the most from this investment requires effort. You can add punctuation by saying ‘full stop’, for instance.  A Quick Reference Card offers a quick overview of voice commands to use when dictating text and when simply using Dragon to navigate the OS. Within Windows 7 and Windows 8 such voice commands are baked in to the accessability options, of course, but if you require voice commands it makes sense to use the same program to dictate as to navigate.

Dragon NaturallySpeaking 12 Premium Expert Verdict »
1.6GHz Intel or AMD processor 3.2GB hard disk space Windows XP, Vista, 7, 8 1GB RAM
sound card supporting 16-bit recording
DVD-ROM drive for installation
  • Ease of Use: We give this item 9 of 10 for ease of use
  • Features: We give this item 8 of 10 for features
  • Value for Money: We give this item 6 of 10 for value for money
  • Performance: We give this item 8 of 10 for performance
  • Overall: We give this item 8 of 10 overall

This is the best speech-recognition dictation software there is. But it is a significant investment, both in terms of time and money. If you already have a recent version of NaturallySpeaking and a decent headset, the upgrade may not be such an attractive proposition. But for those looking for a tool via which they can dictate rather than type, this is the tool for you.

There are currently no price comparisons for this product.
  • Nuance Dragon NaturallySpeaking 11 Premium

    It’s taken two years for Nuance to update its market-leading voice recognition software Dragon NaturallySpeaking. We liked version 10, so what’s new in Dragon NaturallySpeaking 11 Premium?

  • Dragon NaturallySpeaking 9.0 Preferred

    Dragon NaturallySpeaking 9.0 Preferred

    Voice-recognition software has come a long way in the past few years.

  • Dragon NaturallySpeaking

    Voice-recognition software has come a long way since the days when it would understand only about one word in 50. Powerful processors and more sophisticated software mean that talking to your PC to create text onscreen is now a reality. And ScanSoft’s Dragon NaturallySpeaking software is among the best-known packages available.

  • Dragon NaturallySpeaking 10 Professional review


via http://www.speechtechnologygroup.com/speech-blog - Tired of typing? The new Dragon NaturallySpeaking might be just what you need… Source  google.com Dragon NaturallySpeaking remains the number one speech-recognition software tool for dictation. But version 12 brings relatively little new to the party. Here’s our Dragon NaturallySpeaking 12 Premium r ...

Wednesday, October 31, 2012

Speech Recognition for the food service industry

Here's another example of a speech recognition implementation in warehouse management systems.

Lucas Systems, a provider of voice-directed warehouse applications, today introduced the next version of Jennifer FoodSelect, a voice-directed solution designed for foodservice and grocery distribution centers.

The latest release of Jennifer FoodSelect adds configurable voice-directed receiving, putaway, returns and replenishment, in addition to enhanced support for GS1 data standards and food traceability using the Engage Management Services Console. Like all Jennifer solutions, Jennifer FoodSelect combines voice direction with speech recognition (using the Serenade Speech Recognition Engine), and, where appropriate, barcode scanning, key entry, and display capabilities provided in industry-standard multi-modal mobile computers. The solution supports GS1 data standards and provides flexible voice- or scan-based data capture for Global Trade Item Numbers (GTIN), lot numbers, date codes, catchweights, and other information using Motorola, Honeywell, or other standard hardware terminals.

The updated product includes the following:

  • Order Selection. Jennifer FoodSelect supports two-stage PIR picking and single and dual-pallet case picking. For foodservice warehouses, it includes on-demand case label printing that eliminates selector idle time, reduces paper handling, and adds points to individual productivity rates.
  • Truck Loading. Jennifer eliminates loading errors, improves efficiency for warehouse workers, drivers, and customers, and enhances worker safety. Standard, voice-enabled HACCP checklists help meet food safety and traceability requirements.
  • QC/Audit. The integrated QC/Audit module allows managers to better prioritize audits and focus scarce QC resources on the orders that need to be checked.
  • Receiving. Associates can use a combination of scan, screen, and voice to compare physical receipts against purchase orders or ASNs. Discrepancies are identified and data capture is immediate.
  • Putaway. The voice system directs workers through the putaway process, capturing and verifying pallet, item, and location information through speech recognition and barcode scanning.
  • Replenishment. This application can be integrated with voice-directed selection so that let-down requests are processed immediately to avoid short shipments.
  • Returns. Processing returns with Jennifer eliminates clerical steps to improve efficiency, minimize data entry errors, and accelerate inventory updates.
  • Engage MSC. Jennifer’s Wweb-based console includes route planning and management, robust productivity and process tracking, and system configuration tools. Engage allows supervisors to manage and coordinate selection, replenishment and other tasks to optimize overall efficiency and maximize throughput.

Upper Lakes Foods, a regional foodservice distributor in Minnesota, is the first customer to install this latest Jennifer voice solution, with immediate, dramatic improvements in selector accuracy and productivity.

“Jennifer FoodSelect addresses the accuracy and productivity challenges of foodservice and grocery DCs, in addition to supporting highly efficient product track and trace capabilities using industry-best speech recognition and barcode scanning,” says Jennifer Lachenman, vice president of product strategy and business alliances at Lucas Systems, in a statement. “From the beginning our goal in introducing Jennifer FoodSelect was to provide large and small foodservice and grocery distributors with the most comprehensive and configurable voice solution possible. This configurable industry solution approach is the best way to deliver a full-featured product while minimizing customization, speeding implementation, and enabling flexibility for the future. The idea is to accelerate deployment without sacrificing flexibility or important features food distributors need to compete.”

via http://www.speechtechnologygroup.com/speech-blog - Here's another example of a speech recognition implementation in warehouse  management  systems. Source  speechtechmag.com Lucas Systems, a provider of voice-directed warehouse applications, today introduced the next version of Jennifer FoodSelect, a voice-directed solution designed for foodservice ...