A prerequisite for this project is to have Docker installed on your machine. This project utilizes a Docker container to have the frontend communicate with the backend. Another added bonus is that with containerization it makes the environment uniform. If you don't have Docker installed- download it here Before building the containers:
cp .env.example .envMy values looked something like this:
MONGODB_URL=mongodb://localhost:27017
MONGODB_DB_NAME=calls_db
OPENAI_API_KEY={OPENAI_API_KEY}
OPENAI_MODEL=gpt-4o-mini
To run the project run:
docker compose up -d --build
# Seed the database
docker compose run --rm api python seed_from_json.pyIf all is successful, then the web application should be available at localhost:8080 along with the API
To approach the problem, I knew that I had to do the obvious things first. Get a simple CRUD application that builds data based on the data's shape in calls.json and a frontend to represent this data. In my opinion, the best tool suited would be using FastAPI or Express as these are lightweight frameworks to build APIS. For the frontend, I used a React flavor named Vite. For a project of small scale, I thought that using Next.js may be overkill and since I wasn't leveraging the Vercel offerings for the backend, Vite seemed like the better choice. As for the database, I decided to use MongoDB as this didn't require a specific schema and allowed me to iterate quicker as I didn't have to think about database migrations in the longer term. For something of scale, we may want ask questions about how frequently fields are going to be added to a model and if the data that we're importing clean? If so, we may want to choose something like a relational database as we could leverage the benefits of SQL better.
Now if that was the only part of the assignment, this would be complete. We'd ingest our data into our database and just view our calls as they are. Instead, came the interesting came where we'd need to utilize an LLM (I chose 4-o-mini as it was cheap and APIs can get expensive) to classify what actions we'd need to take for a specific call. The best tool in my toolbelt that I could think of would be to have the LLM do sentiment analysis on the data's transcript to give a proposed action, whether or not the action should be human or automated, sentiment, along with the reasoning behind the result. Sentiment analysis is a good choice here as it would primarily be used to predict the emotion of a caller based on their transcript. For new calls, we would add this logic in as a processing step and new calls would show what the LLM would predict. This logic would also be separated away from the logic used in our API directly to separate the concerns of an API and be resuable for something like our ingestion script.
As mentioned previously, I asked the LLM to do sentiment analysis on a given call's transcript, goal, and duration. In the beginning, this received mixed results. For calls that required payments, if a user didn't want to pay it was considered as no escalation necessary and be automated. This meant that the transcript needed to meet the goal or target that they were supposed to. In order to do that- I added a definition of success for each goal so that the LLM had good guidelines of what actions it should provide a user with. I also added more to my prompt to bias certain responses for the human vs automation classification as different emotions would be better for a human agent vs an automated agent.
For calls that went to voicemail or broke up and were otherwise ambiguous, I prompted the LLM to have the proposed action be a follow up call as the goals would not have been known to be met. For calls that had incorrect info or otherwise bad, data I also included rules into the prompt to bias the human agent.
AI was utilized to generate the project's code. The design of the system was done by myself and code tweaks were made to the front and back end of the code in order to get specific parts of the project working correctly. Verification of the generated code was done by human input(myself), as that would be the way to see if the code that was generated met what I saw fit for the design of a system like this.
- Model Used to Classify - I used a cheap model to do my classification and analysis. A better model may be able
- Design - Looks pretty generic, but overall is clear to a user what they'd be doing on the page
- NoSQL DB vs Relational - I touched on it briefly, but this is a bigger trade off. As mentioned earlier there are factors that go into this, but the project could have used a PostgreSQL database instead of MongoDB. I opted not to out of convenience and the iteration that I got back without having to think of the relations each part of the call would have. In a larger system though, we may have wanted to link the actions away from the calls in order to keep the calls separated from what we intend to do with the call. For larger sets of data as well, we'd have to be getting information we wouldn't need instead of storing it in separate tables that are related to each other.
Expecting the action to route to anything. The proposed actions would need specific code actions and a better API for it to do automated tasks.