Jay Myers, SR Software Engineer
June 11, 2024
The thousand injuries of Fortunato I had borne as I best could, but when he ventured upon insult, I vowed revenge…
Written in 1846, The Cask of Amontillado by Edgar Allen Poe is a short story about a man wronged. Montresor plots his revenge on Fortunato by luring him into his family vault with the promise of a rare wine, which Fortunato cannot resist.
…we will go back; your health is precious … and I cannot be responsible…
At several moments on the journey, Fortunato seems to be shown hints of the plan, though he is unable to piece them together. Had Fortunato been able to pick up on these clues, perhaps his fate would have been different. Alas, though the hints become more and more blatant as the journey progresses, he is ultimately unable to see them even at their most prominent.
…“And the motto?”
“Nemo me impune lacessit.”*
Fortunato’s fate is sealed behind a brick wall within the Montresor family vault, and with it, a warning to us all – never insult a Montresor, and pay attention!
Cuing, the potential for one test item to indicate information about another, is one of several ways in which items can be considered enemies of one another. Those who are counting on us to administer effective, fair, and accurate examinations are owed more attention than poor Fortunato was able to provide for himself, and so it falls to us to take on that responsibility.
But how can we accomplish this? Over the years here at Strasz, I’ve been fortunate enough to witness several different strategies and see them change. I’d like to share my experience with you now. Chances are you’re already familiar with some of these techniques, but with any luck, we’ll all learn a little something from this exercise.
*Latin for ‘No one provokes me with impunity’
Manual Review
Good old-fashioned manual review – built on the hard work and experience of SMEs and Test Developers alike, its strengths and weaknesses are born from the same source – the humans behind it. It’s unquestionably true that the proper application of experience separates effective tests from the rest, and building a mechanism in which participants can function efficiently is key. But even the best of us are not immune to the flaws of the human mind. It’s important to acknowledge these shortcomings to try and supplement our efforts with tools and techniques that can improve our results.
The Invisible Gorilla
If you’re unfamiliar with this particular phrase or the experiment behind it, take a peek at the footage here. Your ability may vary, given that I’ve already spoiled a piece of it with the header of this section, but when it was originally conducted, about half of the participants did not notice the gorilla at all. When they’re intently paying attention to one thing, human beings often become somewhat blind to the things around them. Perhaps you’ve noticed this in some aspect of your life, even if you didn’t miss the gorilla.
Asking participants to look for similar items in an enemy analysis may blind them to items that clue one another, or they may miss similarities in distractors across items. Multitasking is difficult for many people, myself included, and asking for an analysis of a list of items, potentially including supplemental passage material, involves many comparisons to keep track of.
Decision Fatigue
When we search for enemy items in a panel of items, we are making a comparison between each item. This means that the first must be compared to every other item, the second item compared to every other item besides the first, as that’s already taken place, and so on. This is a concept in mathematics denoted by the following formula (called n choose k)

Where n is the number of items and k is the number you’re “choosing” – in this case, two as we are comparing each item to each other item once. This means that in a list of 250 items, you’re dealing with a staggering 31,125 individual comparisons. That’s an awful lot of choices to make, especially if you consider that you’re comparing not only similarity but also looking for clues that indicate information between items, comparing distractors, and potentially comparing supplemental information like passages for each item as well.
Research suggests that the more decisions we make in a day, the less able we are to make additional decisions – eventually, we become too drained. In addition to the myriad decisions we all make every day, such as what to wear and what to eat, adding a considerable number of decisions to someone’s workload will drain them. Eventually, you’ll end up with diminished returns and poorly made decisions.
Tools & Tech
So, how do we overcome these obstacles? We develop and use tools to help us, just like we’ve always done. One iteration of these tools I’ve seen in action before the current trend toward AI is a combination of Natural Language Processing and Cosine Similarity Index calculations to compare similarities between items. Then, of course, there’s the ability of relatively newly available Large Language Models to help us in this process, which opens up a whole new avenue of exploration.
In the next post, I’d like to talk about the history of these tools here at Strasz, give you a little insight into how they work, and share some of our future ambitions.
About the Author
Jay Myers is a Senior Software Developer who has spent more than a decade developing and refining testing and testing-related products here at Strasz. From authoring to delivery, standard setting, item analysis, and test prep, he’s been involved in a wide variety of areas supporting test providers and certification bodies. Prior to working at Strasz, he was involved in the Electronic Medical Record space on a product which, statistically speaking, if you’ve been to a hospital in the US has probably been used to treat you. Graduating from Penn State with a degree in Computer Science, he has been interested in solving problems from a young age and his work at Strasz enables him to do that every day. When he’s not working, he spends his time enjoying nature, baking, reading, creative writing, and consuming long form critical analysis of his favorite video games and media.