Quick Guide: Steps To Perform Text Data Cleaning in Python

avcontentteam 12 Jul, 2020 • < 1 min read

Introduction

Twitter has become an inevitable channel for brand management. It has compelled brands to become more responsive to their customers. On the other hand, the damage it would cause can’t be undone. The 140 character tweets has now become a powerful tool for customers / users to directly convey messages to brands.

For companies, these tweets carry a lot of information like sentiment, engagement, reviews and features of its products and what not. However, mining these tweets isn’t easy. Why? Because, before you mine this data, you need to perform a lot of cleaning. These tweets, once extracted can come with unwanted html characters, bad grammar and poor spellings – making the mining very difficult.

Below is the infographic, which displays the steps of cleaning this data related to tweets before mining them. While the example in use is of Twitter, you can of course apply these methods to any text mining problem. We’ve used Python to execute these cleaning steps.

Download the PDF Version of this infographic and refer the python codes to perform Text Mining and follow your ‘Next Steps…’ -> Download Here

To view the complete article on effective steps to perform data cleaning using python -> visit here

If you like what you just read & want to continue your analytics learning, subscribe to our emails, follow us on twitter or like our facebook page.

advanced data cleaning data cleaning data science infographic infographics python text mining

avcontentteam 12 Jul 2020

Beginner Business Analytics Infographic Infographics NLP

Responses From Readers

Indraneel Pise 30 Jun, 2015

Stemming is also an important step in text mining. You could include that too.

Rohit 30 Jun, 2015

remember to deal with character encodings.

Rohit Shetty 30 Jun, 2015

Thou r awwssomee

Teresa Rothaar 30 Jun, 2015

This is great. I'm doing a project on social media data mining for my capstone class for my MS in MIS, and this is very helpful. Thanks so much!

Whaly 20 Jul, 2015

This is awesome, can you check the urls in the last step? They seem not working :)

mzkarim 30 Oct, 2015

Great resource. Thanks.

Asmaa M. Elmohamady 12 May, 2018

i need help in slang lookup

Write for us

Write, captivate, and earn accolades and rewards for your work

Reach a Global Audience
Get Expert Feedback
Build Your Brand & Audience

Cash In on Your Knowledge
Join a Thriving Community
Level Up Your Data Science Game

imag

Rahul Shah

Sion Chakrabarti

CHIRAG GOYAL

Barney Darlington

Suvojit Hore

Arnab Mondal

Prateek Majumder

7