Python Data Analysis · Lesson 1 of 10

Free preview

The analysis workflow, and setting up pandas

Set up pandas, NumPy and Matplotlib in a virtual environment and take a first, honest look at a messy dataset.

time
55 min
on completion
+110 XP

Why this matters

A manager forwards you a file called tickets.csv and asks a simple-sounding question: "how are we doing on support?" The file has about five thousand rows. Before you can answer anything you have to open it, work out what each column means, notice that some rows are blank, that the same word is spelled three different ways, and that a few tickets appear to have been closed before they were opened. Only then can you count, group and chart.

That is the real shape of data analysis. Surveys of working analysts put the majority of their time on finding, loading and cleaning data, not on the analysis everyone imagines. This course teaches the whole loop on a realistic, deliberately messy dataset, using pandas, the standard Python library for tabular data. Pandas is what analysts, data engineers and scientists reach for when the data is bigger than a spreadsheet is comfortable with, or when the work has to be repeatable rather than clicked by hand.

The workflow you will use in every lesson is: load the data, inspect it honestly, clean it, transform it into the columns you need, summarise and communicate the result, and do all of it in a script you can re-run. This first lesson sets up the tools and takes that honest first look.

Concepts

pandas, NumPy and Matplotlib. *pandas* gives you two objects. A Series is one labelled column of data; a DataFrame is a whole table, a dictionary of Series that share an index. *NumPy* is the numeric engine underneath pandas; you will use it directly for missing values (numpy.nan) and for fast conditional columns. *Matplotlib* draws charts, which you will save to image files rather than pop up on a screen, because the lab has no desktop.

A virtual environment. Ubuntu manages its own system Python and will refuse to let pip install into it (you will meet the externally-managed-environment error if you try). The supported way to install your own packages is a virtual environment: a private copy of Python whose pip installs land in a folder you own. You created these in the Python course; here every project gets one.

The dataset. You will work with helpdesk, generated for this course: about 5,000 support tickets from a fictional managed-service provider, plus the customers, agents and service plans they relate to. It is synthetic — no real person or company — but it carries the flaws real exports have, so the cleaning you learn is the cleaning you will actually do at work.

Reproducibility. A result you cannot reproduce is not a result. Throughout, you keep your work in .py scripts in a project directory. Anyone can re-run a script and get the same answer; that is the difference between an analysis and a screenshot.

Guided exercise

  1. Create the project and a virtual environment on linux01, and install the three libraries. Pandas pulls in NumPy automatically, but naming it is clearer:

``bash mkdir -p ~/analysis && cd ~/analysis python3 -m venv .venv source .venv/bin/activate python -m pip install pandas numpy matplotlib

What you will see, and why it looks alarming. The lab machine has no route to the internet on purpose. It carries its own local copy of these packages (a *wheelhouse*) at /opt/ultiblob/wheelhouse, and /etc/pip.conf points pip at it. So pip tries the internet first for each of the twelve packages it installs, fails to resolve the name five times, and then installs from the local copy. The command prints about 75 lines, 60 of them WARNING, and ends like this:

``text Looking in links: /opt/ultiblob/wheelhouse WARNING: Retrying (Retry(total=4, ...)) after connection broken by 'NewConnectionError(... Temporary failure in name resolution')': /simple/pandas/ ... Processing /opt/ultiblob/wheelhouse/pandas-3.0.5-cp312-...-x86_64.whl Installing collected packages: six, pyparsing, pillow, packaging, numpy, kiwisolver, fonttools, cycler, python-dateutil, contourpy, pandas, matplotlib Successfully installed contourpy-1.4.0 cycler-0.12.1 fonttools-4.65.0 kiwisolver-1.5.1 matplotlib-3.11.2 numpy-2.5.3 packaging-26.3 pandas-3.0.5 pillow-12.3.0 pyparsing-3.3.2 python-dateutil-2.9.0.post0 six-1.17.0

The warnings are noise; the last line is the result. Installing from a vetted local wheelhouse instead of the public index is normal practice on production and regulated networks, so this is worth recognising now: read the last line, not the loudest one.

  1. Put the data in place. The lab provides it at ~/analysis/data/; confirm it is there:

``bash ls ~/analysis/data

``text agents.csv customers.csv plans.json tickets.csv

If the folder is empty (for example on your own machine), generate it from the course dataset: python3 /path/to/dataset/generate.py -o data.

  1. Write ~/analysis/01_first_look.py. Four methods answer "what did I just load?": .shape (rows, columns), .head() (the first rows), .dtypes (the type pandas chose per column) and .info() (a summary with non-null counts and memory).

```python import pandas as pd pd.set_option("display.width", 120) pd.set_option("display.max_columns", 20)

tickets = pd.read_csv("data/tickets.csv")

print("shape=", tickets.shape) print(tickets.head()) print(tickets.dtypes) tickets.info(memory_usage=True) ```

  1. Run it from the project directory with the venv active:

``bash python 01_first_look.py

The shape is the first thing to check, and it already tells a story — 5,030 rows:

``text shape= (5030, 14)

The info() summary is the honest part. Read the non-null counts, not just the column names. Your own output lists all fourteen columns; the rows that matter are these four:

``text RangeIndex: 5030 entries, 0 to 5029 Data columns (total 14 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 2 closed_at 4777 non-null str 3 customer_id 4970 non-null float64 7 agent_id 4798 non-null float64 12 satisfaction_score 2659 non-null float64 dtypes: float64(4), int64(2), str(8) memory usage: 550.3 KB

Three things should already bother you. closed_at has only 4,777 values out of 5,030, so some tickets are still open. customer_id is a float64 even though an id should be a whole number — that happens because some are blank. And satisfaction_score is missing on nearly half the rows. You have not cleaned anything yet; you have just looked, which is exactly the point of the first step.

Troubleshooting

Dozens of `WARNING: Retrying ... Temporary failure in name resolution` lines from `pip` → the lab has no internet by design and pip is falling back to the local wheelhouse → nothing is wrong; check the last line says Successfully installed ... pandas-3.0.5 .... The one message that *is* a real failure looks different: ERROR: Could not find a version that satisfies the requirement ... (naming the package) means it is not in the wheelhouse, and the lessons only use packages that are.

`error: externally-managed-environment` when you run pip → you are installing into the system Python → create and activate the virtual environment first (python3 -m venv .venv then source .venv/bin/activate); your prompt should now start with (.venv).

`ModuleNotFoundError: No module named 'pandas'` → the venv is not active, or you ran python3 (the system one) instead of the venv's python → re-run source .venv/bin/activate, or call .venv/bin/python 01_first_look.py explicitly.

`FileNotFoundError: data/tickets.csv` → you are not in ~/analysis, or the data was not seeded → cd ~/analysis and check ls data; regenerate with the dataset script if needed.

Your output looks different from a blog post you found → this course targets pandas 3.0, which changed several defaults (string columns show as str, not object; other differences appear later). Match the version in course.yaml, not an older tutorial.

`SyntaxWarning` or odd wrapping in `head()` → widen the display with the pd.set_option lines shown, or select fewer columns; it is only a display setting, not a data problem.

Check your understanding

  1. Why is customer_id loaded as float64 rather than an integer type?
  2. What does .info() tell you that .head() does not?
Answers

1. The column has blank cells. NumPy's integer type cannot hold a missing value, so pandas widens the whole column to float64, where missing becomes NaN. You will fix this in the cleaning lesson. 2. .head() shows a few example rows; .info() shows every column's non-null count and dtype at once, so it reveals missing data and type surprises that a handful of rows would hide.

Summary and next step

  • Analysis is a loop: load, inspect, clean, transform, summarise, communicate — and keep it in a script.
  • Install pandas, NumPy and Matplotlib into a virtual environment; never into the system Python.
  • shape, head, dtypes and info are your first four looks, and info's non-null counts already

expose missing data and type surprises.

Next, in *Loading data*, you load CSV and JSON properly and learn why the dtypes come out the way they do.

References

  • pandas documentation (3.0), Getting started and Intro to data structures — https://pandas.pydata.org/docs/ (accessed 2026-09-13)
  • pandas, What's new in 3.0.0 — https://pandas.pydata.org/docs/whatsnew/v3.0.0.html (accessed 2026-09-13)
  • Python Packaging User Guide, Installing packages using pip and virtual environments — https://packaging.python.org/en/latest/guides/installing-using-pip-and-virtual-environments/ (accessed 2026-09-13)

Sign in to record your progress

Signing in saves eligible progress. It does not enroll you or include a lab; review the career path for access terms.

Sign in