The manual exists and nobody opens it
Most companies handle an unfamiliar task the same way. The person doing the work asks someone who has done it before. The knowledge exists, and in most cases it has also been written down, in a manual or in a recording of a training session. The difficulty is not that the work was never documented. It is that documentation cannot be consulted in the seconds an employee has while a customer waits, so the question goes to a colleague instead.
A company uploads the files it already holds. The system reads each one and expresses the work it describes as a numbered sequence of steps.
What the employee gets back is the answer itself, in steps, and not a recording with an instruction to find the relevant part. Underneath that answer sits the source it came from, the page of a document or the second of a recording, so the employee can check it.
What follows is the whole path, from the recording to the answer at the counter, one step at a time, together with the research behind two of its design decisions and the three things it still does badly.
What the phone call really costs
Asking a colleague is not free. The cost falls mainly on the person interrupted, and it has been measured.
When knowledge is reachable only through a person, every question carries three costs. The employee who asks learns nothing durable and will ask again. The employee who answers becomes the person everybody interrupts and cannot stop being it. The customer waits while the two of them resolve it.
The second of those has been measured. Gloria Mark's team at UC Irvine timed 24 office workers for more than 700 hours: 57% of their work segments were interrupted, and getting back to an interrupted task took 25 minutes and 26 seconds. Sharing a room made it worse, 63% interrupted against 49% for those working apart, which is the case a shop floor is in.
They were watching analysts and developers, not a counter, and nobody has run that study behind one. We are borrowing the mechanism, not the setting.
When only a person can answer
- The employee asks a colleague or telephones the manager
- The answer depends on who is available
- Two people stop what they were doing
- Nothing is recorded for the next occurrence
- The manual and the recording remain unopened
When the files can answer
- The employee asks in their own words and reads the steps
- The answer comes from the same file for every location
- Nobody else is interrupted
- The source opens at the page or the second it came from
- Material already recorded is finally used
What Hellomatik is
Hellomatik is an assistant that employees write questions to. It answers using the files their company has uploaded, not general knowledge about the world.
The agent is the assistant employees interact with. It answers only from material the company has uploaded, and every answer identifies which file it used.
This is the distinction that matters in practice. A general assistant knows how the software works. This one knows how your company uses it, because your files are the only thing it reads. It does not know Excel. It knows how your team builds its reports.
The knowledge base is where those files are stored: manuals in PDF, recorded training sessions, price lists, internal documents. Each file is uploaded once and remains available.
A procedure is what the system writes after reading one of those files. It expresses the work described in that file as numbered steps in the order they occur, including the points at which the work branches depending on the case. A return within thirty days and a return after thirty days are different sequences, and a procedure represents them as such.
We are currently deploying this in large companies, where the requirement is usually consistency rather than training. Above a certain number of locations, each one drifts toward its own version of a process, and that drift is normally discovered through a customer complaint.
Step 1. Somebody records the job
One employee performs a task while describing it aloud. That recording is the input to everything that follows.
Take the one this post follows: processing a return after the thirty day window has closed. Somebody who does it every week records themselves doing it, with a phone propped against a box. No script, no editing, no second attempt.
Two conditions matter. The audio has to be intelligible, and the work has to be described aloud rather than only performed. "Now I open the order using the customer's email address" can become a step. A silent sequence of clicks cannot, because there is nothing for the system to read.
Step 2. The file is read, cut and timestamped
The system splits the file into short passages so that a single one can be cited, and for video it also writes down the second at which every word was said.
An MP4 works. So do PDFs, spreadsheets and text documents. Nothing has to be recorded again.
The return recording is not stored as one lump. It becomes a few hundred short passages, each keeping a record of where in the file it came from. That is what lets an answer point at one passage, and the passage point back at the page or the minute.
Video gets one more thing, and everything later depends on it. Every word is transcribed with the second it was spoken and every speaker is labelled, so a training session with questions from the room does not become one undifferentiated block of text.
By the end of this step the recording is no longer a recording: it is a list of sentences, and every one of them knows its own second.
Step 3. The recording becomes a procedure
The system reads the transcript and writes the work out as a start, the actions in order, the points where it branches, and an end.
Forty minutes of recording are not forty minutes of knowledge. From the return recording, what comes out is four steps and one decision: open the order at 0:00, check the purchase date at 1:14, ask the supervisor at 2:48, close and notify the customer at 4:23.
The branch is the part a summary would lose. Inside thirty days and past thirty days are different jobs. The procedure keeps them apart, and writes the steps they share once, after they rejoin.
We cut the recording into steps. Reviewing the segmenting principle, Richard Mayer counts ten experiments out of ten in which people learn a process better from segments they advance themselves than from the same material played straight through. The conditions where that matters most describe the counter exactly: complex material, explained fast, to somebody who has not done it before.
Steps have to name the real button and the real field. One that reads "enter the details" is rejected and rewritten. That is exactly the instruction that sends the employee back to the phone.
Step 4. You correct it first
The procedure sits next to the file it came from, and you can read it before anybody on the floor does.
Review is optional and we recommend it. It is where you find out that the recording skipped the case of a gift receipt. Or that "new customer" was read more broadly than you meant it.
Correcting it means correcting the source file and generating the procedure again, which is a few minutes of your time against a whole floor having already read the wrong version.
Step 5. The answer, opened at 2:48
The employee writes a question. The answer arrives with the file and the exact place inside it that it came from.
Somebody three weeks into the job types "how do I process a return after 30 days". What comes back is the answer, in steps: past thirty days it needs a supervisor, send the order number in the store channel, wait for the approval.
Underneath the answer sits the proof. The recording, opened at 2:48, with the four steps listed alongside, so that moving between the steps moves the recording. The second is there so the employee can check the answer, not so they have to sit through it.
And the employee can keep asking. "And if the supervisor is not in?" is a different branch of the same procedure. It is answered the same way, from the same recording.
From the procedure itself there is a fuller version. It goes one step at a time, asks at the branch which case applies, and waits for the employee to choose. It remembers where they left off and carries on the next day.
The agent does not memorise any of this. It looks the answer up in your files on every question. And it reports where it found it, which is what makes a single step checkable.
How a procedure is stored
As a flow: one start, a series of actions, branches that carry their own paths, and one end. Timestamps are attached to the actions and to the branches, never to the start or the end, because those two are not moments in the recording.
How the steps become chapters of the recording
Every step carrying a timestamp becomes a chapter that can be jumped to, in the same order as the procedure. Where a step arrives without a description, the chapter states that in plain terms, and the entry is never left empty.
The same recording, another voice
Any video in the knowledge base can be given a second voice track in another language. The picture and the timings are unchanged.
This is a branch off the main path, and for a workforce that does not share a language it decides one thing: whether the recording circulates at all.
What is produced is not a subtitle. It is the same recording with a different voice over it, which matters because an employee working at a counter cannot read subtitles and perform the task at the same time.
Two properties are worth stating explicitly. The voice is generated automatically and is not reviewed by a person. And only the language being watched is downloaded, never the full set, which is what keeps it usable on a tablet in a store.
The expert leaves, the recording stays
An employee on a shop floor performs many separate tasks. Each is recorded once, and the resulting material serves both the employee learning and the employee checking.
Opening and closing the till, processing a return, requesting a transfer from another location, dressing a seasonal window, enrolling a customer in a loyalty scheme. In the roles we have recorded so far, around twenty of them.
A single procedure then serves two different needs, and this is the aspect of its use that surprised us most. It serves learning, when a new employee works through a procedure from beginning to end. And it serves reference, months later, when an infrequent case has been forgotten. The material is identical in both cases. What differs is whether the employee follows the whole sequence or requests one step.
We leave the choice to the employee, because guidance that carries a beginner turns redundant for somebody who already holds the process in their head. Kalyuga and his co-authors named it the expertise reversal effect: techniques that work with inexperienced learners lose their effect on experienced ones, and can even backfire.
There is a second effect, and it is the one companies notice later. The person who recorded the procedure is usually the one who has been there longest. When that person leaves, what they knew normally leaves with them. Here the recording stays, and so does the way they did the work, alongside everything else the company has written down.
This has a direct consequence for the objective usually described as closing the gap between experienced and new staff. That gap is not closed by putting both through the same course. It is closed by leaving the material in segments and letting each employee take only the portion they lack.
What this does not do well yet
These are the limits, and we would sooner state them than have you find them.
The jump to a specific second can fail. Where the system cannot determine with confidence which part of a recording an answer was drawn from, the video opens from the beginning. The file and the procedure are still correct. What is lost is the shortcut.
A silent recording produces nothing. The system works from what is said out loud, so a screen capture with no narration leaves you a file the agent can play and no procedure at all. If a library is mostly silent screen recordings, the useful part of this starts by recording them again with somebody talking.
Dubbing is generated automatically and the voice is not reviewed by a person. It is adequate for an internal procedure and it is not adequate for customer facing material. We would rather state that than have it discovered in production.
Frequently asked questions
- Can the agent invent steps that are not in the file?
- It records only actions that the person in the source material actually performs, and excludes the surrounding explanation. That is a rule the system is required to follow rather than a mathematical guarantee, which is why the procedure can be read and corrected before any employee queries the agent.
- Does it work with PDFs, or only with video?
- With both, and with word processor documents, plain text and web pages. Spreadsheets are the exception: a table of rates has no sequence of actions in it, so it contributes to answers but does not produce a procedure. What only video provides is a timestamp attached to each step.
- How long does the process take?
- The time is spent uploading and processing the file, and it depends on the length of the recording. The procedure is then generated from the document in seconds. Review takes as long as the company decides to spend on it.
- What happens when a process changes?
- The new version is uploaded and its procedure generated. Replacing the old file stops the old version being used in answers, and from that moment every location is reading the new one. There is nothing to resend and nobody to notify. Retaining both files leaves both citable, which is rarely the intended outcome, so replacement is the normal course.
- Does it work with a recording in which several people speak?
- Yes. The transcription separates the speakers, so a session with one presenter and questions from an audience does not become a single undifferentiated block of text.
The employee never watched the recording. It answered them anyway, at 2:48, and answered them again when they asked what happens if the supervisor is out.
References
- 1.Mark, G., Gonzalez, V. M. and Harris, J. (2005). No Task Left Behind? Examining the Nature of Fragmented Work. CHI 2005, Portland, 321-330. Observation of 24 information workers over more than 700 hours. ↩
- 2.Mayer, R. E. and Pilegard, C. (2014). Principles for Managing Essential Processing in Multimedia Learning: Segmenting, Pre-training and Modality Principles. In R. E. Mayer (ed.), The Cambridge Handbook of Multimedia Learning, 2nd ed., 316-344. Cambridge University Press. ↩
- 3.Kalyuga, S., Ayres, P., Chandler, P. and Sweller, J. (2003). The Expertise Reversal Effect. Educational Psychologist, 38(1), 23-31. ↩
- 4.Everything stated here about how Hellomatik works comes from our own implementation, read off the code in July 2026. ↩