Skip to content

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Wikipedia-Text-Classification-Project

Project Description: Text Classification for Medical/Non-Medical Attribution

Overview: Develops a system attributing English texts to medical/non-medical classes via Wikipedia analysis. Utilizes NLTK, SpaCy (Python), OpenNLP, Naive Bayes, Logistic Regression, and processing techniques (Bag of Words, Stop Word List, Stemming, Lemmatization). Leverages Wikipedia API for annotated texts.

Implementation Options:

NLTK/SpaCy Implementation:

  • Develops models using NLTK/SpaCy.
  • Applies Naive Bayes, Logistic Regression.
  • Uses Bag of Words, Stop Word List, Stemming, Lemmatization.
  • Utilizes pre-annotated Wikipedia texts.

Python Implementation

  • Implements the system in Python (OpenNLP).
  • Explores Naive Bayes, Logistic Regression.
  • Applies Stop Word List, Stemming, Lemmatization.
  • Accesses Wikipedia API for annotated texts.

Project Structure:

Directories:

  • data: Holds training/testing datasets, including pre-annotated Wikipedia texts.
  • src: Source code for the text classification system.
  • models: Stores trained models.
  • docs: Project documentation, API guidelines, and the project report.

Usage Guidelines:

  • Offers clear setup and run instructions.
  • Provides guidelines for Wikipedia API access and netiquette adherence.
  • Documents preprocessing, model training, and evaluation processes.
  • Lists dependencies and system requirements.

Contribution Guidelines:

  • Forks the repository.
  • Creates branches for features/bug fixes.
  • Ensures code aligns with the style guide.
  • Writes clear commit messages.
  • Submits pull requests for review.

References:

GitHub Project Link

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages