Forum mining challenge – Get the right questions!

Kunal Jain Last Updated : 24 Apr, 2015

3 min read

Analytics community is relatively small but very vocal on the web world. We subscribe to various Linkedin groups, Facebook groups and other websites to be aware of all that is going on in the industry. For instance, Stack overflow has more than 8 million questions asked till date. (Source: http://data.stackexchange.com/) .

The problem statement we are going to provide in this article is a good combination of data extraction and unsupervised modeling algorithms and will not only make you understand how to fetch data from forum websites but also make you understand what is the analytics industry wondering about.

Background

Analytics Vidhya was built with a vision of creating a strong analytics community ready to share knowledge and best practices. With such aspirations, Kunal and Tavish wish to answer questions on which people in the industry are either confused or need guidance. Now, answering more than a million questions posted till date on stack overflow is next to impossible.

It’s your turn to help Analytics Vidhya recognize the popular topics of discussions & have oodles of fun during the process.

Problem statement

Find the most popular forum questions (related to data-science) across stack overflow network. Now, as mentioned – this can be huge! So, here are some guidelines to identify relevant questions for your task:

Questions posted with various data related tags (e.g. data-mining, data.frame, data-science etc.) across the network
Questions on site dedicated to data science (e.g. datascience.stackexchange.com or stats.stackexchange.com)

Please note that these are guidelines and not a definition to use. Like all open problems, there might be pockets, these guidelines are missing.

Also, in order to keep the analysis relevant, you should include questions posted after Jan 2012.

What Data you need to use for categorization?

All the questions posted after 1st January 2012 (included) on stack overflow network can be used to train the model. You can use all the fields which can be obtained directly from the API of stack overflow. This can include topic of discussion, number of views, date of publish, author of the question etc.

Help : Text Mining tools on R and Clustering algorithms will come very handy to build the solution.

What is the evaluation metric?

The total popularity (views) of the top fifty questions found will be used as the evaluation metric. For instance, if the first question category found is “merging tables in SAS” and the number of views for the following questions are:

“How to merge two tables in SAS?” : 50

“Joining tables in SAS?” : 100

” Problem while finding linked rows between 2 datasets in SAS” : 200

The total number of views counted for the first question category is the total of all the linked questions (i.e. 350 in this case).

Submission Format :

We expect following three parts for the submission :

1. Codes built to do this analysis

2. Top 50 question category found

3. Mapping file, which should include all the linked questions to the question categories. We will do a search on these questions to calculate the total number of views. Note that all the questions put in the same category should have the same answer.

All submissions should reach us latest by 15th October 2014 on [email protected]

End Notes:

The aim of this challenge is to foster analytical thinking in our reader’s mind and have some fun with practical machine learning / analytics challenges!
We will give the winner of this challenge a chance to blog about his solution on Analytics Vidhya. Of course, he takes away all the visibility, which comes on the platform!
Last but not the least, the entire story presented before is hypothetical. It was created with the sole aim to create this challenge.

If you like what you just read & want to continue your analytics learning, subscribe to our emails, follow us on twitter or like our facebook page.

Kunal Jain

Kunal Jain is the Founder and CEO of Analytics Vidhya, one of the world's leading communities of Al professionals. With over 17 years of experience in the field, Kunal has been instrumental in shaping the global Al landscape. His expertise spans diverse markets, from developed economies like the UK to emerging ones like India, where he has successfully led and delivered complex data-driven solutions. As a recognized thought leader, Kunal has empowered countless individuals to realize their Al ambitions through his visionary approach to Al education and community building. Before founding Analytics Vidhya, Kunal earned both his undergraduate and postgraduate degrees from IIT Bombay and held key roles at Capital One and Aviva Life Insurance across multiple geographies. His passion lies at the intersection of analytics, Al, and fostering a thriving community of data science professionals.

Free Courses

4.7

Generative AI - A Way of Life

Explore Generative AI for beginners: create text and images, use top AI tools, learn practical skills, and ethics.

4.5

Getting Started with Large Language Models

Master Large Language Models (LLMs) with this course, offering clear guidance in NLP and model training made simple.

4.6

Building LLM Applications using Prompt Engineering

This free course guides you on building LLM apps, mastering prompt engineering, and developing chatbots with enterprise data.

4.6

Improving Real World RAG Systems: Key Challenges & Practical Solutions

Explore practical solutions, advanced retrieval strategies, and agentic RAG systems to improve context, relevance, and accuracy in AI-driven applications.

4.7

Microsoft Excel: Formulas & Functions

Master MS Excel for data analysis with key formulas, functions, and LookUp tools in this comprehensive course.

Reading list

Forum mining challenge – Get the right questions!

Background

Problem statement

What Data you need to use for categorization?

What is the evaluation metric?

Submission Format :

End Notes:

If you like what you just read & want to continue your analytics learning, subscribe to our emails, follow us on twitter or like our facebook page.

Login to continue reading and enjoy expert-curated content.

Free Courses

Generative AI - A Way of Life

Getting Started with Large Language Models

Building LLM Applications using Prompt Engineering

Improving Real World RAG Systems: Key Challenges & Practical Solutions

Microsoft Excel: Formulas & Functions

Recommended Articles

Responses From Readers

Become an Author

Flagship Programs

Free Courses

Popular Categories

Generative AI Tools and Techniques

Popular GenAI Models

AI Development Frameworks

Data Science Tools and Techniques

Reading list

Basics of Machine Learning

Machine Learning Lifecycle

Importance of Stats and EDA

Understanding Data

Probability

Exploring Continuous Variable

Exploring Categorical Variables

Missing Values and Outliers

Central Limit theorem

Bivariate Analysis Introduction

Continuous - Continuous Variables

Continuous Categorical

Categorical Categorical

Multivariate Analysis

Different tasks in Machine Learning

Build Your First Predictive Model

Evaluation Metrics

Preprocessing Data

Linear Models

KNN

Selecting the Right Model

Feature Selection Techniques

Decision Tree

Feature Engineering

Naive Bayes

Multiclass and Multilabel

Basics of Ensemble Techniques

Advance Ensemble Techniques

Hyperparameter Tuning

Support Vector Machine

Advance Dimensionality Reduction

Unsupervised Machine Learning Methods

Recommendation Engines

Improving ML models

Working with Large Datasets

Interpretability of Machine Learning Models

Automated Machine Learning

Model Deployment

Deploying ML Models

Embedded Devices

Forum mining challenge – Get the right questions!

Background

Problem statement

What Data you need to use for categorization?

What is the evaluation metric?

Submission Format :

End Notes:

If you like what you just read & want to continue your analytics learning, subscribe to our emails, follow us on twitter or like our facebook page.

Login to continue reading and enjoy expert-curated content.

Free Courses

Generative AI - A Way of Life

Getting Started with Large Language Models

Building LLM Applications using Prompt Engineering

Improving Real World RAG Systems: Key Challenges & Practical Solutions

Microsoft Excel: Formulas & Functions

Recommended Articles

Responses From Readers

Become an Author

Flagship Programs

Free Courses

Popular Categories

Generative AI Tools and Techniques

Popular GenAI Models

AI Development Frameworks

Data Science Tools and Techniques