VFourCuts is a photo booth application written in Python that performs real-time gesture recognition and automatically captures photos when a V-sign is detected. This program uses OpenCV and Tensorflow for gesture recognition and captures and composes images after a countdown.
- The script uses the MediaPipe library for hand landmark prediction and recognizes the V-sign(peace) gesture in real-time camera frames.
- A pre-trained TensorFlow/Keras model for hand gesture recognition is loaded using
load_model. The model is designed to predict gestures based on hand landmarks.
- The script captures video frames from the default camera (
cv2.VideoCapture(0)). - A countdown is initiated when the V-sign(peace) gesture is detected, and images are captured every 5 seconds for a total of 4 repetitions.
- Hand landmarks are extracted, and the V-sign(peace) gesture is predicted for each detected hand.
- Landmarks are visualized on the frames using
mpDraw.draw_landmarks. - Images are combined into a single image after a 5-second countdown.
Supported Python version: 3.11
- Install Python: Download and install Python from the official Python website Make sure you are using a supported version of Python.
- Download the sources: Run
git clone https://github.com/heoheosilsil/vFourCuts.gitor download the zip archive of this project. - Install the dependencies: In the project directory, use the command
pip install -r requirements.txtto install the dependencies.
- Run the
main.pyfile. - Once the program starts, gesture detection will start.
- When a V-Sign is detected, a 5 second countdown will start and a photo will be taken.
- A total of 4 photos will be composited and saved in a file named
result_vfourcuts.png.
When the program is running, you will see:
- A countdown timer
- The current number of detected V-Signs
The project uses OpenCV for image processing and Tensorflow for gesture recognition. A detailed list of dependencies and their versions can be found here.
This program uses pre-trained datasets from the following website for machine learning gesture recognition.
- The technique requires nearly perfect top-down views for accurate gesture recognition.
If you would like to contribute to this project, please open a pull request or issue. All contributions are welcome!
