Showing posts with label VUI Research. Show all posts
Showing posts with label VUI Research. Show all posts

Thursday, June 12, 2008

Design and measure the "itties"

How are you designing and measuring your product or service's major "itties?" If you're a usability engineer then you focus on usability - making the product easy to use for consumers. But what about the other "itties?" Here's my short list of the major itties.
  • Usability
  • Utility
  • Identifiability
  • Trustability

Everyone gets usability, right? UEs design and test, and design some more until customers can use the product (it's usually a product, since we don't usability test services much, yet) easily. Utility, we mostly get that. The product is supposed to meet a customer's need. In the worst case, ensuring utility means walking down a list of features the product is supposed to contain. If we're really conscientious, we create a contextual task analysis and design a "whole product" solution, in which the product either supplies a complete solution to a need, or the product complements are easily available to the customer.

Identifiability is the product's distinctiveness, its uniqueness, the design edge that sets it apart from its competition. Think iPod. Product marketers get this, I think, but there are an awful lot of "me too" product and e-commerce web sites out there. We recognize identifiability when we see it, but practice it too infrequently.

Trustability isn't obvious at all, but when I talk to user experience designers about it, something just clicks. Lots of people have had the experience of finding something for sale on a e-commerce site but were unwilling to purchase from the site. Why? The site was easy to use, it was distinctive, it offered the product we wanted and good delivery terms, but we just couldn't click the "purchase" button. We went elsewhere. This is a classic trustability issue. We didn't trust the site, or the company, or something else, and someone lost a sale.

I've done a fair amount of research on trustability. It's a great topic. As a VUI designer I've found that the quality of the system's voice has an aspect of trustability.

All of the "itties" that I mentioned are important, but my observation has been that people involved in product development tend to focus their efforts on one or two. Designing and measuring the full customer experience requires that product designers address all of the itties.

Wednesday, May 7, 2008

Design anti-patterns: What's one more option?

You have just finished a solid draft of an IVR call routing application, and you're pretty happy with. The menus follow conventional wisdom for numbers and types of options, and the design has done well in early usability testing. You're looking forward to sign-off on the design, and one of the business analysts pipes up, "Hey, we need to add one more option to the main menu. What's one more option?"

"What's one more of anything," like records in a database or fields on a form, usually isn't a big deal. Computers are supposed to be able to iterate any number of things with a minimum of additional effort, so the same principle should apply to IVR menu options. That's the logic, anyway. Unfortunately, the same logic applied to IVR menus is a disaster, because it's the callers who have trouble sorting through and remembering and verbally reproducing multiple menu options.

The logic of "What's one more option" is very seductive, and can lead to main menus with ten or more options. Big main menus (or "Maim menus," for what they do to callers) just scream "I don't know what I'm doing! Hit the zero key now!"

IVR menus structures are still actively researched. For an excellent recent study on IVR menu options, see "A comparison of broad versus deep auditory menu structures," Commarford et al., (2008), Human Factors, Vol. 50, pp. 77-89.

Friday, October 12, 2007

Low bar vs. high bar usability tests

Here are three usability test situations that I've encountered in the past.

  • An interaction designer looks at the graphical page headers she's been sent by a graphics designer and something looks wrong: the words on the images seem blurry and hard to read. She tries to enhance the images in Photoshop and asks me for me opinion of the enhanced images. We quickly design a comparative usability test of the images. The original images (A) are placed next to the enhanced (B) images, and three questions appear below each image: (1) on a 1-5 scale which image appears sharper (1=image A is much sharper, 5=image B is much sharper)? (2) Which image is easier to read? and (3) Which do you prefer? The test materials are put in a Word document and sent to 12 office co-workers. The tests are returned and tabulated in less than three hours-the enhanced images are judged sharper, easier to read, and are preferred by a clear margin. The results are sent to the business lead with the recommendation to consider using the enhanced images until further testing is completed. The business lead who employed the graphic designer rejects the recommendation, saying that the data weren't valid because the test participants weren't real customers.


  • I'm designing a main menu for a touchtone IVR. I'm not familiar with the terms used by the business for the menu options, but the client assures me that callers will understand the options, even though callers have no more knowledge about that business than I do. I type up some alternative wordings for the menu options then mail them out to some coworkers with the question below each set of options: please list the types of services or products you would expect to find associated with each option. I collect my responses for the next two days, and the results aren't encouraging. I show my results to the client and suggest a different strategy for naming the menu options, with follow up testing on the proposed menus. The test results are rejected because the methodology is too dissimilar to an actual IVR experience.


  • I'm designing a menu for a speech IVR and I'm not sure the menu options the business wants are going to be discriminable to the speech recognizer. I ask that a one-menu prototype be created so I can test the discriminability of the options. The request is refused by a manager because, I'm informed, I'm not a real customer so it's not a valid test.

In each case the designs of the tests that I've described barely qualify as usability tests. They don't follow the procedures that you might find in classic texts on usability testing like the Handbook of Usability Testing or A Practical Guide to Usability Testing. The first example looks more like the procedure for an eye exam: "which do you prefer, A or B? A or B?" The second example looks like a poorly conducted card sort. The third looks more like a QA procedure. So what's an experimental psychologist like me doing running these poorly designed tests?

I call these sorts of usability tests low bar usability tests. They are simple, easy to design and execute usability tests that answer the question, "is it OK to proceed with this design or idea?" Since many interface designs take a good deal of time and effort to complete, and usability tests themselves can take a good deal of time to perform, a low bar usability test can tell the designer if he or she is on the right track. If the design passes the low bar test, design continues, and the next usability test, the high bar usability test, is run according to the canonical texts. If the design fails the simple low bar test, there's no point in taking the design in that particular direction and running a high bar test. It's time to step back and re-design.

To take one example, if I test the goodness of some menu options in an IVR I'm designing and I can't get the recognizer to reliably distinguish my utterances, should I continue with this design and get data from real customers? Of course not. I've worked with IVRs for years and I'm very good at getting them to do what I want them to do. If I can't use the menu no one else will be able to. If the menu passes this low bar test, does it show that the menu is ready for prime time? Of course not. I still need to run a high bar usability test. But at least I can move forward with the knowledge that the menu has a chance of working.

You should add low bar usability testing to your discount usability bag of tricks, remembering to explain the rationale for running a less-than-perfect test procedure to those who learned usability testing by the book.

Wednesday, October 3, 2007

Upcoming: Conversations 2007 conference

My company is sending me to the Conversations 2007 conference Oct. 21-24 in Boca Raton, Fla. I'm always happy to attend conferences, meet people, and learn about what's going on the industry. The conference is hosted by Nuance. Company-sponsored conferences have a different flavor than academic conferences (e.g., HFES). There's a lot more selling done, a lot less data presented, but there are more opportunities to meet business contacts. I'll have a full report when I get back.

Saturday, August 25, 2007

Customer preference for live CSR?

Customers prefer to speak to a real person instead of a speech IVR: we in the speech industry have heard it said a million times, and never with any data to back it up. Now, there's a survey that claims Australians prefer to speak to IVRs than to CSRs in overseas call centers.

Now, the first thing I want to know when I read about a survey is the methodology used, the questions asked, and the see the real data. This CNET article that I linked to doesn't give specifics, and the published report is very expensive, so I can't vouch for the quality of the research. I'm quite sure that the survey respondents were thinking of efficient, well-designed IVRs, since there's almost nothing so aggravating than a malfunctioning IVR. But, if the generalization that a group of English-speaking callers prefer speech IVRs to overseas CSRs, then there exists data that partially refutes the argument that customers always prefer to speak to a live CSR.

Personally, I always take a shot at using self service on the phone, since I never know how long I'll be waiting in queue before talking to a person. Maybe it's just professional curiosity, but I like seeing how different IVRs operate. And if the IVR can't give me what I want, I haven't lost anything. I'm assuming that a lot of people feel the same way, but I don't have any data. The CNET article is a pointer to some real research that should be replicated in the US and elsewhere.

Thursday, August 16, 2007

Speech recognition for handhelds

When Google gets into something, it does it in a big way. I've blogged about the well done 800 GOOG 411 speech-driven search service before. It made me wonder what else Google was developing with speech.

Then I saw this article about speech recognition and GPS for handheld devices. The article rightly points out that manual input is difficult to design and implement on small handheld devices, and speech (once you get the recognition working correctly) is a natural candidate for input. Hey, there's no training involved - we already know how to talk into small devices, right?

The article states "Google research director, Peter Norvig, has indicated that Google is currently spending more on speech and translation than any other area." The article goes on to suggest that Google will produce a competitor to the iPhone that incorporates both speech and GPS. This makes sense. Google took a big step away from desktop-based search with GOOG 411. Enabling location-based search on an easy to use handset is a next, very large, logical step. I'm guessing that Google will partner with a company that produces handsets to implement their ideas. It will be interesting to see what they develop.

Monday, July 30, 2007

Catching up on the Dragon Systems founders

Jim and Janet Baker founded Dragon Systems back in 1982. The company produced Dragon Dictate, the first speech recognition system for PCs, and Dragon Naturally Speaking. They sold the company in 2000. They were among the first to use Hidden Markov Models to represent spoken language.

I read an interesting news item on Jim Baker recently. He's moving from Carnegie Mellon to Johns Hopkins University to work on a new research project for the Defense Department and the National Security Agency. The NSA wants to be able to conduct surveillance on millions of phone conversations, and today's speech recognition software isn't up to the task. Thus, they've funded one of the founders of the field to get speech recognition unstuck.

I'm conflicted about this. I admire Baker for the pioneering work he did on speech recognition. I understand that his initial research was funded by the DoD through ARPA, and much good came of that research eventually. Baker is good enough to be able to move the engineering of speech recognition forward. However, I don't trust the motives of the NSA, and if this is made to work properly we'll have on our hands another Big Brother-styled technology available to the government for eavesdropping on US citizens.

Maybe we'll get lucky. Maybe the funding will run out, the program will be seen as an expensive failure, and the NSA will say "so long and good luck." Then Baker will release some really groundbreaking research that will be of benefit to everyone working in speech. That's what I'd like think.

Tuesday, June 26, 2007

Awash in real interaction data

In the past I've lamented the lack of published VUI user data. If you go to the big design conferences (HFES, CHI, UPA, etc.) you'll find lots of data from studies on web navigation and mobile device use, but very little on speech interfaces. I know this because when I present a study at these conferences the speech session (if there is one) is very small. The science and practice is built on good data, and there is much to go on right now.

There are, in fact, lots of data. Nearly all implementations go through at least one tuning cycle in which caller utterances are recorded and transcribed and matched to the speech recognizer's responses. These data are then analyzed for things such as grammar coverage and the appropriateness of the prompts and so on. Each tuning exercise generates a LOT of data. Unfortunately, the data are usually locked up by the company conducting the tuning exercise, never to see the light of day. There are legitimate reasons why companies don't share their data. Publication of the data is of little value to the company that produces it, and competitors could take advantage of it without responding in kind. Even if scrubbed, data could be used to identify the company or business units for which the data were produced. So I can see why companies are disinclined to distribute their data.

Still, I wonder what sort of intervention could be put in place to motivate companies to share their tuning data for the purpose of improving the best practices in the industry. If anyone has ideas, let me know.