print("Hello World")College of Computing and Informatics · Study summary
DS 231: Introduction to Data Science Programming
In [1]:course = { "code": "DS 231", "title": "Introduction to Data Science Programming", "part_1": "Data science concepts (Modules 2-5)", "part_2": "Python for data science (Modules 6-12)", }
Every module of the course in plain language: what each idea means, why it matters, and the details most likely to show up in a quiz. Modules 2–5 cover the data science concepts, and Modules 6–12 teach Python.
Orientation
Big picture. The course has two halves. First you learn how data is collected, organised, explored and used. Then you learn Python to get insights out of that data.
Part 1: Fundamentals
Standards and practices for collecting, organizing, managing, exploring and using data.
Part 2: Python
The techniques needed to get insights from data using Python.
Why this course matters
Many important decisions made by people and by society are (or should be) data driven. The primary goal is experience in manipulating, analyzing and presenting data. The second goal is being able to analyze a problem and run the analysis with Python.
Textbooks
| Book | Used for |
|---|---|
| Data Science, 3rd Edition (2021), Lillian Pierson | Modules 2–6 (concepts, plus open data in Chapter 19) |
| Python All-in-One, 2nd Edition (2021), John C. Shovic and Alan Simpson | Modules 6–12 (Python) |
Course learning outcomes
- Identify and describe the methods and techniques commonly used in data science.
- Apply data preprocessing tools.
- Develop Python code elements.
- Analyze and evaluate results of data science solutions.
- Show how data analysis, modeling, machine learning and statistical computing work together.
Quick check
What are the two parts of the course?
Fundamentals of working with data, and getting insights from data using Python.
What is the primary goal of the course?
Gaining experience in manipulating, analyzing and presenting data.
Wrapping Your Head around Data Science
Big picture. Data science turns raw data into useful decisions. Doing it well takes three skills together: math and statistics, coding, and knowledge of the subject area.
Data is everywhere
Data comes from every computer, phone, camera and sensor. It is created whenever you post, save a file, take a picture or search for the nearest ice cream shop. Data engineers find ways to capture and condense huge volumes of data, and data scientists find valuable, actionable insights in it. Using those insights is “like being able to see in the dark”.
Data science vs. data engineering
Data science
The computational science of extracting meaningful insights from raw data and communicating them to create value.
Data engineering
An engineering field that builds and maintains systems to handle large volumes, varieties and velocities of data.
Often confused.The two terms are misused as if they were the same thing. Engineers build the pipes; scientists analyse what flows through them.
The three varieties of data
| Type | What it means | Example |
|---|---|---|
| Structured | Stored and processed in a traditional relational database (RDBMS), in clean rows and columns. | A MySQL database |
| Unstructured | Generated by human activity; does not fit a structured database format. | Email documents |
| Semistructured | Does not fit a structured database, but is organised by tags that give it order and hierarchy. | XML and JSON files |
Who can use data science?
It used to be only large tech companies. Today organizations of all sizes are in a “sink-or-swim”, data-driven competitive environment, so data know-how is a core skill in almost every line of business. Anyone with a bit of training can use data insights to improve their life, career or business.
The pieces of the data science puzzle
A real data scientist needs math and statistics, coding skills and subject matter expertise (SME). These six components appear in every data science role:
1. Collecting, querying and consuming data
Data engineers capture and collate big data, but data scientists also query data constantly, from one database or several. Four universal ways to work with data:
| Format / tool | Notes |
|---|---|
| CSV (comma-separated values) | Accepted by almost every desktop and web analysis application. |
| Script files | .py or .ipynb for Python, .r for R. |
| Application | Excel, good for quick analyses of small to medium datasets. |
| Web programming | D3.js (data-driven documents), a JavaScript library for custom web visualizations. |
2. Mathematical modeling
Math uses deterministic methods to describe the world with numbers. Data scientists use it for predictive forecasting, decision modeling and hypothesis testing.
3. Statistical methods
Statistics helps you understand significance, validate hypotheses, simulate scenarios and forecast. The basics to learn: linear regression, logistic regression, naive Bayes classification and time series analysis.
4. Coding
5. Applying data science to a subject area
Statisticians say “it’s just statistics with a new name”. The main difference is the need for subject matter expertise. Example: a clinical informatics scientist combines healthcare expertise with data science to build personalized treatment plans.
6. Communicating insights
If you cannot communicate, all the insight in the world helps nobody. Scientists must explain insights so staff understand them, make clear visualizations and written narratives, and be creative and pragmatic.
Career paths
Data implementer
Builds data and AI solutions. Detail-oriented. Math and coding are the power.
Data leader
Leads teams and stakeholders to successful data solutions. Loves the outcomes data science makes possible and collaborates a lot.
Data entrepreneur
Builds businesses by delivering data science services and products. Craves the creative freedom of being a founder.
Quick check
Define data science and data engineering in one line each.
Data science extracts meaningful insights from raw data and communicates them. Data engineering builds and maintains systems that handle large volumes, varieties and velocities of data.
Give an example of each data variety.
Structured: MySQL database. Unstructured: email documents. Semistructured: XML or JSON files.
What separates data science from traditional statistics?
The need for subject matter expertise, plus the use of programming languages and approaches from mathematics.
Which career path suits someone who wants to found a business?
Data entrepreneur.
Tapping into Critical Aspects of Data Engineering
Big picture. Big data is data too big, too fast or too messy for ordinary databases. Handling it needs special storage and processing tools, in the cloud or on your own servers.
What is big data? The three Vs
Big data exceeds the processing capacity of conventional database systems because it is too big, moves too fast, or lacks the structure traditional databases need. Using it requires big data storage and processing, such as a Hadoop cluster. If volume, velocity or variety forces you to adopt such a solution, you have a big data problem.
Volume
How much. The lower limit starts at about 1 terabyte and has no upper limit. Many tiny transactions in many formats.
Velocity
How fast. Data volume per unit of time, from about 30 KB/s up to 30 GB/s. Often created by automated processes.
Variety
How different. Structured plus unstructured and semistructured data: JSON, XML, graph data, social media, weblogs, click-streams.
Remember.Hadoop is powerful at one thing: batch-processing and storing large volumes of data. It reduces big data into smaller datasets that scientists can analyze.
Important data sources
Humans, machines and sensors generate data constantly. Typical sources:
Three roles that are often confused
| Role | What they do | Typical skills |
|---|---|---|
| Data scientist | Discovers knowledge through data analysis. Makes sense of the data. | Math and statistics, programming, domain expertise |
| Machine learning engineer | A software engineer who knows enough data science to deploy models inside applications (bring models to production). | More computer science and software development than a typical data scientist |
| Data engineer | Builds and maintains systems that store, migrate and process data at high volume, velocity and variety. | Java, C++, Scala, Python; Hadoop MapReduce, Spark; RDBMSs; MPP platforms |
Hiring rule.Hire a data engineer to store, migrate and process data; a data scientist to make sense of it; and a machine learning engineer to put models into production.
Data science can, for example, optimize energy usage with machine learning, predict unknown contaminant levels from sparse environmental data, and detect fraud and theft by spotting anomalies.
Storing and processing data in the cloud
Cloud storage offers faster time-to-market, more flexibility and security.
Serverless computing
Your model runs directly in its container in the cloud. The provider adjusts the infrastructure, so data scientists spend less time preparing it.
Kubernetes
Open-source software that manages, orchestrates and coordinates containerized applications across clusters. Works on-premise, in the cloud or in a hybrid cloud.
NoSQL databases
A traditional RDBMS only handles clean rows and columns queried with SQL. It cannot handle unstructured or semistructured data, or big data volume and velocity. NoSQL databases are non-relational, distributed systems built for big data. They are schema-free, can run on-premise or in the cloud, and let you query data without SQL.
Storing big data on-premise: Hadoop
The Hadoop platform was designed for large-scale processing, storage and management. It is open source and has four main parts:
- MapReduce is a parallel distributed framework that processes huge volumes in batch. It converts raw data into sets of tuples, then combines and reduces them into smaller sets of tuples.
- HDFS stores data on clusters of commodity hardware. Three key features: blocks, redundancy, fault tolerance.
- MPP (massively parallel processing) is an alternative to MapReduce. MPP runs on costly custom hardware; MapReduce runs on inexpensive commodity servers.
Processing big data in real time
A real-time framework processes data as it streams into the system. It either improves overall time efficiency or uses new querying methods for real-time queries. In-memory processing works inside the computer’s memory without writing results to disk, so it is much faster, but it cannot process much data per interval.
Quick check
What are the three Vs?
Volume, Velocity and Variety.
What does each part of Hadoop do?
HDFS stores, MapReduce batch-processes, Spark processes in real time, YARN manages resources.
Why use NoSQL instead of an RDBMS for big data?
An RDBMS handles only clean relational rows and columns. NoSQL is non-relational and distributed, so it copes with unstructured and semistructured data at big data volume and velocity.
MPP vs MapReduce?
Both do distributed processing. MPP uses costly custom hardware; MapReduce uses inexpensive commodity servers.
Trade-off of in-memory computing?
Very fast results, but it cannot process much data per processing interval.
Machine Learning: Using a Machine to Learn from Data
Big picture. Machine learning applies algorithmic models to data over and over so the computer discovers hidden patterns and uses them to make predictions.
What is machine learning?
It is also called algorithmic learning. Use cases include real-time internet advertising, spam filtering, recommendation engines, natural language processing and sentiment analysis, and automatic facial recognition.
Not the same thing.Machine learning, data science and AI are separate. Machine learning is a practice within data science, and data science is more than machine learning.
The three steps of the machine learning process
- Acquire data
- Preprocess it
- Feature selection
- Split into training and test sets
- Model experimentation
- Training
- Building
- Testing
- Model deployment
- Prediction
Rule of thumb for splitting data: randomly sample two-thirds of the dataset to train the model, and use the remaining one-third as test data to check the accuracy of its predictions.
Machine learning vocabulary
ML borrows words from both statistics and computer science. Different names, same idea:
| Machine learning term | Also called |
|---|---|
| Instance | Row (data table), observation (statistics), data point, case |
| Feature | Column or field (data table), variable (statistics), independent variable (IV) in regression |
| Target variable | Predictant, dependent variable (DV) |
Learning styles
Supervised
Input data has labeled features. The model learns from known labels to predict labels for new data. Use it when you have labeled historical values that predict future events.
Example: logistic regression
Unsupervised
Accepts unlabeled data and groups observations into categories based on similarities in the input features.
Examples: PCA, k-means clustering, SVD
Reinforcement
Behavior-based, like how humans and animals learn. The model gets rewards and learns to maximize the sum of its rewards by adapting its decisions.
The book names supervised, unsupervised and semisupervised as the three main styles (semisupervised is “an up and coming star”), and the slides then explain reinforcement learning as a third behavior-based model.
Choosing algorithms by function
Algorithm families are grouped by what they do. Examples from the book’s figure:
| Family | Examples |
|---|---|
| Regression | Linear regression, OLS, logistic regression |
| Regularization | Ridge, LASSO, Elastic Net, LARS |
| Bayesian | Naive Bayes, Bayesian Belief Network |
| Decision tree | CART, C4.5, CHAID, Decision Stump |
| Dimensionality reduction | PCA, PLSR, MDS, Sammon Mapping |
| Instance based | k-nearest neighbour (kNN), Self-Organizing Map |
| Clustering | k-means, k-medians, Expectation Maximization, hierarchical clustering |
| Ensemble | Random Forest, Boosting, Bagging, AdaBoost |
| Neural networks | Perceptron, Back-Propagation, Hopfield Network |
| Deep learning | Convolutional Neural Network, Deep Belief Network, Stacked Auto-Encoders |
| Rule system | Cubist, OneR, ZeroR, RIPPER |
Deep learning in real life: Gmail’s Smart Reply (the three one-line auto-replies) and Facebook’s DeepFace (automatic tag suggestions for people in photos).
Apache Spark
An in-memory distributed computing application for deploying machine learning algorithms on big data sources in near-real-time, producing analytics from streaming data.
Quick check
Name the three steps of the machine learning process.
Setup, learning, application.
How should data be split for training and testing?
Randomly use two-thirds for training and keep one-third for testing.
Supervised vs unsupervised?
Supervised uses labeled data to predict labels (e.g. logistic regression). Unsupervised uses unlabeled data and finds groups by similarity (e.g. k-means, PCA).
What is a feature? What is a target variable?
A feature is a column/variable (the independent variable). The target is the dependent variable you want to predict.
Probability and Statistical Modeling
Big picture. Statistics, probability and math are the tools data scientists use to understand data and predict what comes next: probability, correlation, dimensionality reduction, decision models, regression, outlier detection and time series.
1. Probability and inferential statistics
A statistic is a result from a mathematical operation on numerical data. It comes in two flavors:
Descriptive statistics
Describe a dataset: its distribution, central tendency (mean, min, max) and dispersion (standard deviation, variance). They show relationships between X and Y but do not claim X causes Y.
Inferential statistics
Take a smaller section of the data (a sample) to deduce things about the larger dataset (the population). Methods like regression do try to predict by studying causation.
Descriptive statistics are also used to detect outliers, plan feature preprocessing, and decide which features to use. Inferential statistics are used when collecting data on the whole population is unaffordable or impossible. For a valid inference, the sample must be representative, and even then it contains random variation called noise.
Probability distributions
Events must be defined as mutually exclusive (only one can occur at a time, like the result of rolling a die). Two rules:
- The probability of any single event is never below 0.0 or above 1.0.
- The probabilities of all events always sum to exactly 1.0.
| Distribution | Kind | Meaning and example |
|---|---|---|
| Discrete | Counted by groupings | Car color: a limited set of values |
| Continuous | Range of values | Car miles per gallon: each car has its own value |
| Normal | Numeric, continuous | Symmetric bell curve; the most likely value is at the top, extremes are less likely |
| Binomial | Numeric, discrete | Number of successes in a number of attempts with only two possible outcomes (binary variables) |
| Categorical | Non-numeric | Categories or ordinal (ordered) variables, such as first, business and economy class |
Conditional probability with Naive Bayes
Naive Bayes predicts the likelihood that an event happens given evidence in your data features (conditional probability). It is especially useful for classifying text, such as spam classification.
2. Quantifying correlation
Correlation is measured by r, between −1 and +1. The closer |r| is to 1, the stronger the correlation. An r close to 0 may mean the variables are independent.
| Method | Use it for | Assumptions |
|---|---|---|
| Pearson’s r | Dependent relationships between continuous variables (the simplest form) | Normally distributed data, continuous numeric variables, linearly related |
| Spearman’s rank | Ordinal variables. Converts variable pairs into ranks. | Variables are ordinal and related nonlinearly |
3. Reducing dimensionality with linear algebra
Arrays and matrices are the main data structures in analytical computing, so you need linear algebra to work with large, multidimensional datasets. The goal of dimension reduction is to compress the data while removing redundant information and noise.
SVD
Singular Value Decomposition. Decomposes data to reduce dimensions. Applied to large, noisy, sparse data to find principal components.
Factor analysis
Filters out redundant information and noise. A variable with more variance carries more information; shared variance means redundancy.
PCA
Principal Component Analysis. Unsupervised. Finds relationships between features and reduces them to non-redundant principal components.
Factor analysis assumptions
- Features are metric, and continuous or ordinal.
- More than 100 observations and at least 5 observations per feature.
- The sample is homogenous.
- Correlation between features is r > 0.3.
PCA vs factor analysis.(1) PCA does not look for an underlying cause of shared variance; it decomposes the data to summarize its most important information in fewer features. (2) With PCA you do not choose the number of components at first: run it, let the results tell you how many to keep, then rerun to extract them.
4. Multiple criteria decision-making (MCDM)
Use MCDM when a decision depends on two or more criteria and it is unclear which has priority. Two assumptions must hold:
- Multiple criteria evaluation: more than one criterion to optimize.
- Zero-sum system: improving one criterion must cost at least one other.
Fuzzy MCDM evaluates suitability on a range of acceptability instead of the crisp 0-or-1 membership of traditional MCDM.
5. Regression methods
Regression came from statistics. It describes and quantifies relationships between variables and shows the strength of correlation. You can predict future values from historical ones, but be careful: regression assumes a cause-and-effect relationship.
Linear regression
Relates a target y to predictor features. With one predictor it is just y = mx + b.
Logistic regression
Estimates values for a categorical target. Also gives the probability of each estimate. The target should be numeric, describing the class.
OLS
Ordinary least squares. Fits a regression line to a dataset. Handy with several independent variables.
Limitations of linear regression
- Numerical variables only, not categorical.
- Missing values cause problems; fix them first.
- Outliers make results inaccurate.
- Assumes a linear relationship between features and target.
- Assumes features are independent of each other.
- Residuals (prediction errors) should be normally distributed.
6. Detecting outliers
Outliers are values significantly different from most points of a variable. Many methods assume there are none, so removal is part of data preparation. Detection can also reveal fraud, equipment failure or cybersecurity attacks.
Univariate (one feature at a time)
Multivariate (two or more together)
7. Time series analysis
A time series is a collection of values of an attribute over time. It is analysed to predict future values from past observations. Univariate time series model changes in a single variable over time.
Four patterns a time series can show:
Quick check
Descriptive vs inferential statistics?
Descriptive statistics describe a dataset (no causation claimed). Inferential statistics use a sample to deduce information about a population and try to predict, studying causation.
Two key rules of probability?
Each probability is between 0 and 1, and all event probabilities sum to exactly 1.
Pearson vs Spearman?
Pearson: continuous, normally distributed, linear. Spearman: ordinal, nonlinear, uses ranks.
What are the two MCDM assumptions?
More than one criterion, and a zero-sum system where improving one criterion sacrifices another.
Name two univariate and two multivariate outlier methods.
Univariate: Tukey outlier labeling, Tukey boxplotting. Multivariate: scatter-plot matrix, DBScan (also boxplotting, PCA).
Open Data Resources and Starting with Python
Big picture. Part 1: where to find free, public data to practise on. Part 2: why Python is the language of data science and how to install the tools.
Part 1: Open data
Open data is data that has been made publicly available and may be used, reused, built on and shared. It belongs to the wider “open movement” (open source software, open hardware, open access journals). Open licenses use copyleft instead of copyright. Governments release open government data to promote transparency and accountability, and volunteers use it to build solutions to social problems.
Be careful.Work labeled “open” may not fit the accepted definition. You are responsible for checking the licensing rights and restrictions of the data you use.
| Source | What to know |
|---|---|
| Data.gov (USA) | Open access to non-classified US government data. Economic, environmental, STEM, quality of life and legal indicators. Over 60 open-source APIs. |
| Canada Open Data | Over 200,000 datasets: environment, citizenship, quality of life. |
| data.gov.uk | Started in 2010 (late to the movement). Environment, government spending, health, education, business. |
| US Census Bureau | Demographics for marketing: age, income, household size, gender or race, education. |
| NASA (data.nasa.gov) | All non-classified project data since 1958. About 4 TB of new earth-science data per day. Astronomy, climate, life sciences, geology, engineering. |
| World Bank | Download datasets or view visualizations; has an Open Data API. Agriculture, economy, environment, financial sector, poverty. |
| Knoema | 500+ databases and 150 million time series from governments, the UN and corporations. |
| Quandl | Toronto-based search engine for numeric data. Links to over 10 million datasets, including 2.1 million UN datasets. |
| Exversion | Like GitHub for data: version control and hosting. All uploads are public. Great for data cleanup. |
| OpenStreetMap | Open, crowd-sourced alternative to Google Maps. Users trace and label routes, for example via phone GPS. |
Saudi Open Data Portal (data.gov.sa)
- A national initiative for a public data hub that supports transparency, e-participation and innovation.
- Publishes datasets from ministries and government agencies in an open format, giving the public one central access point to find, download and use them.
- Benefits: understand how agencies work, evaluate their performance, make informed decisions about policy, and build research, reports, web and mobile apps.
- Facts: 6,544 datasets (August 2022), 147 publishers, use cases available, you can request a dataset. The open data license lets users distribute, transform and build on the data, with the source credited.
Part 2: Why Python?
In 2017 Python became the most popular language in the world (IEEE Spectrum). Three reasons:
- It is relatively easy to learn.
- Everything you need is free.
- It has more ready-made tools for data science, machine learning, AI and robotics than most languages.
Choosing the right Python
Go to python.org. It recommends the most current stable version. Use that one.
Tools for success
- An editor lets you type code; an interpreter runs it. You need both.
- Anaconda is a complete Python development environment with a friendly graphical interface, often called a data science platform because many of its packages are data-science oriented. It also comes with VS Code.
- Install: go to anaconda.com/download, download the free version, follow the prompts, and choose to install Microsoft VS Code when asked.
- Anaconda Navigator lets you move between features and launch what you want.
- Open VS Code from Anaconda with its Launch button. Check the Extensions icon (puzzle piece): you should see Anaconda Extension Pack, Python and YAML.
- Jupyter Notebook is another popular tool for writing Python. It supports Julia, Python and R, is free with Anaconda, and is often used to share code online.
Quick check
What is open data?
Data made publicly available that may be used, reused, built on and shared. Open licenses use copyleft instead of copyright.
Which portal publishes Saudi government datasets?
The Saudi Open Data Portal, data.gov.sa.
Editor vs interpreter?
The editor is where you type code; the interpreter runs it.
What does Anaconda include?
A full Python environment with data science packages, Anaconda Navigator, VS Code and Jupyter Notebook.
Interactive Mode, Getting Help, and Writing Apps
Big picture. Python can be used in two ways: typing commands one at a time in the interpreter (interactive mode), or writing a .py file or a Jupyter notebook and running it.
Interactive mode in the VS Code terminal
- Open Anaconda Navigator and Launch VS Code.
- If the Terminal pane is hidden, choose View → Terminal.
- Check your version at the operating system prompt (space before the first hyphen, no other spaces):
python --version
You should see something like Python 3.x.x. Now start the interpreter:
python >>>
The >>> prompt means you are inside the Python interpreter. Type a command and press Enter. A spelling mistake gives an error message.
Built-in help
>>> help() help> keywords
help>means you left the interpreter (>>>) and are in the help area. Type a module, keyword or topic name to get help.keywordslists the words with special meaning in Python.- Leave help with
q(quit) or Ctrl+Z, then leave Python withexit(). - Online help: YouTube for videos, Stack Overflow for questions, plus search engines.
Workspace, folder and your first file
- A VS Code workspace is your development environment: the Python interpreter you use plus any extensions you add. Store it anywhere.
- Create a folder for all your Python code and associate it with the workspace, so you always use the right interpreter and settings.
- Each Python file is a plain text file with the
.pyextension. The VS Code editor is a code editor.
Your first program, Hello World:
Hello World
- Saving: VS Code does not save automatically. Save often, or turn on File → Auto Save (a check mark means it is on).
- Running: right-click the file name and choose Run Python File in Terminal. The terminal shows the output and a new prompt.
Simple debugging
VS Code shows errors before you run: red file and folder names, a red error count, a red circled X total at the bottom left, and a wavy red underline on the bad code.
Python is case-sensitive.PRINT in capitals is an error. The correct command is print.
VS Code also has a built-in debugger for more complex programs.
Jupyter Notebook
- Notebooks are saved with the
.ipynbextension. Save with File → Save and Checkpoint. - A notebook is made of cells. A Code cell holds Python; the output appears below it.
- Run a cell with Alt+Enter (Windows), Option+Enter (Mac) or the Run button. Only the cell with the cursor runs; the double-triangle icon runs all cells.
- Markdown cells hold formatted text, pictures and video. Markdown is like a greatly simplified HTML. Insert a cell with Insert → Insert Cell Below.
- To reopen, launch Jupyter Notebook from Anaconda and click the file.
Quick check
How do you check the Python version?
Type python --version at the terminal prompt.
What do >>> and help> mean?
>>> means you are in the Python interpreter. help> means you are in interactive help.
How do you run a .py file in VS Code?
Right-click the file name and choose Run Python File in Terminal.
How do you run a Jupyter cell?
Alt+Enter (Windows), Option+Enter (Mac), or click Run.
Python Elements and Syntax
Big picture. Python is designed for humans to read. It uses objects, requires indentation to mark code blocks, and gets extra powers from importable modules.
The Zen of Python
Many languages focus on what the computer does. Python is built around how humans think, work and communicate. Python code should be more human-readable than machine-readable.
Object-oriented programming (OOP)
OOP mimics the real world: it consists of objects with properties and methods (actions). Think of a car:
Class
The object creator, like a car factory that can produce many kinds of cars.
Object
One thing made from a class: one particular car. Exists only inside the computer.
Properties and methods
Properties differ between cars. Methods (steer, speed up, brake) are shared by all.
Python is very much object-oriented, but first learn the core language so you know how to use other people’s objects.
Why indentation counts, big time
| JavaScript and similar | Python | |
|---|---|---|
| Marks a block with | Parentheses, curly braces, semicolons | Indentation only |
| Indentation is | Optional (just for readability) | Required and changes how the code runs |
Remember.In Python the indentation is the structure. No indentation, or wrong indentation, means broken code.
Using Python modules
Python has a simple, clean core language, plus hundreds of free modules written and tested by other people for specific jobs (science, AI, dates and times). You import a module and use it as its documentation says. Newer versions may exist, but if your version works you need not upgrade.
import random
answer = random.randint(1, 8) # random whole number from 1 to 8Core Python has no random number generator built in, so import random brings it in.
Import syntax
import modulename [as alias]
- Written in lowercase exactly as shown (case-sensitive).
- Italic words are placeholders you replace with your own.
- Anything in [square brackets] is optional, and you never type the brackets.
- Put import statements first in the file.
Aliases
An alias is a nickname after the word as. It saves typing when module names are long.
import random as rnd
answer = rnd.randint(1, 8)Quick check
What is the main philosophy of the Zen of Python?
Code should be geared to how humans think and communicate: more readable by humans than by machines.
Class vs object?
A class is the creator (like a car factory); an object is one thing it produces (one car).
How does Python mark a block of code?
By indentation, which is required.
What does import random as rnd do?
Imports the random module and gives it the alias rnd, so you write rnd.randint(1, 8).
Building Your First Python Application
Big picture. The building blocks of Python code: comments, data types, operators and variables, put together following correct syntax.
Comments
A comment is text that does nothing when the program runs. It explains the code to teammates and to your future self. Two ways:
# This is a Python comment (starts with a pound sign)
"""This is a multiline comment.
It is sometimes called a docstring.
It starts and ends with three double quotation marks."""Data types
Computers can do arithmetic on numbers, not on words.
Numbers
Integer: whole number, e.g. 0, -1.
Float: has a decimal point, e.g. 1.1.
Complex: ends with j (imaginary part).
Strings
Text. Always in quotation marks, single or double. Digits inside quotes are still a string.
Booleans
Only True or False. No quotes, and the capital first letter is required.
What is a valid number?
It starts with a digit, a decimal point or a minus sign, and has at most one decimal point, with no letters, spaces, $ or commas.
| Good | Bad (why) |
|---|---|
1, 1.1, 1234567.89, -2, .99 | $1.99 (has $), 12,345.67 (comma), 1101 3232 (space), 91740-3384 (hyphen), (267)555-1234 (parentheses and hyphen), 127.0.0.1 (more than one decimal point) |
Quotes in strings
"Mary's dog said Woof" # apostrophe inside: use double quotes
'The dog of Mary said "Woof".' # double quotes inside: use single quotes
x = True # Boolean, no quotes, capital TOperators
| Arithmetic | Meaning | Example |
|---|---|---|
+ | Addition | 1 + 1 = 2 |
- | Subtraction | 10 - 1 = 9 |
* | Multiplication | 3 * 5 = 15 |
/ | Division | 10 / 5 = 2 |
% | Modulus (remainder after division) | 11 % 5 = 1 |
** | Exponent | 3 ** 2 = 9 |
// | Floor division | 11 // 5 = 2 |
| Comparison | Meaning |
|---|---|
< | Less than |
<= | Less than or equal to |
> | Greater than |
>= | Greater than or equal to |
== | Equal to |
!= | Not equal to |
is | Object identity |
is not | Negated object identity |
| Boolean | Example | True when |
|---|---|---|
or | x or y | Either x or y is True |
and | x and y | Both x and y are True |
not | not x | x is not True |
Variables
A variable is a placeholder for information that may change. The = sign is the assignment operator: it assigns the value on the right to the variable on the left.
x = 10 # store 10 in the variable x
quantity = 3
unit_price = 6.5
extended_price = quantity * unit_price
print(extended_price)19.5
Rules for variable names
- Must start with a letter or an underscore (
_). - After the first character you can use letters, numbers or underscores.
- Case-sensitive:
Priceandpriceare different. - No quotation marks inside the name.
- PEP 8 style: lowercase letters with underscores between words (
unit_price). Use meaningful names.
Syntax
Syntax is the proper order and format of code, as important as in human language. If you break it, Python shows a SyntaxError. A line of code ends with a line break or a semicolon, and both of these run the same:
first_name = "Alan"
last_name = "Simpson"
print(first_name, last_name)
first_name = "Alan"; last_name = "Simpson"; print(first_name, last_name)The exercises in this module (type, save, run, change, save again, run again) are what you do in any software development, so practise them until they feel natural.
Quick check
Two ways to write a comment?
Start with #, or enclose text in triple quotation marks (docstring).
Name Python’s basic data types.
Numbers (integer, float, complex), strings, Booleans.
What is 11 % 5? And 11 // 5?
11 % 5 = 1 (remainder). 11 // 5 = 2 (floor division).
Is 2var a valid variable name?
No. A variable name must start with a letter or an underscore.
What is = called?
The assignment operator. Use == to test equality.
Working with Numbers (Part 1)
Big picture. Python’s built-in and math functions do calculations for you, f-strings control how numbers are displayed, and Python can write numbers in binary, octal and hex.
Calling functions
variablename = functionname(param[, param])
Most functions return a value, so you store it in a variable. The values inside the parentheses are parameters (arguments). abs() takes one number; round() takes a number and then the number of decimal places.
| Built-in function | Purpose |
|---|---|
abs(x) | Absolute value of x |
bin(x) | x as a binary string |
float(x) | Convert a string or number to a float |
hex(x) | x as hexadecimal, prefixed 0x |
int(x) | Convert to integer by truncating (not rounding) the decimals |
max(x, y, z, ...) | The largest argument |
min(x, y, z, ...) | The smallest argument |
oct(x) | x as octal, prefixed 0o |
round(x, y) | Round x to y decimal places |
str(x) | Convert the number to a string |
type(x) | The data type of x |
format(x, y) | Older formatting; replaced by f-strings |
z = -999.9999
print(abs(z))
print(int(abs(z)))
print(round(3.14159265, 4))
print(type(3.14), type(128), type(str(5)))999.9999 999 3.1416 <class 'float'> <class 'int'> <class 'str'>
The math module
More functions live in the math module. Import it first, then write the module name, a dot, and the function name.
import math
print(math.sqrt(81))
print(math.factorial(7))
print(math.floor(-23234.5454))
print(math.radians(45))9.0 5040 -23235 0.7853981633974483
| Function | Purpose |
|---|---|
math.sqrt(x) | Square root |
math.pow(x, y) | x raised to the power y |
math.factorial(x) | Factorial of x |
math.ceil(x) / math.floor(x) | Smallest integer ≥ x / largest integer ≤ x |
math.log(x, y) / math.log2(x) | Logarithm of x to base y / base 2 |
math.exp(x) | e raised to the power x |
math.sin / cos / tan(x) | Trigonometry, in radians (acos, atan, atan2 also exist) |
math.radians(x) / math.degrees(x) | Convert degrees to radians / radians to degrees |
math.isnan(x) | True if x is not a number |
math.pi, math.e, math.tau | Constants: 3.14159…, 2.71828…, 6.28318… |
Formatting numbers with f-strings
Put a lowercase f (or F) right before the opening quote. The text inside is the literal part (shown exactly). Anything in curly braces {} is the expression part: a variable or a calculation that gets filled in.
quantity = 3
unit_price = 499.90
print(f"Subtotal: ${quantity * unit_price:,.2f}")Subtotal: $1,499.70
| Format | Meaning |
|---|---|
:, | Comma in the thousands place |
.2f | Fixed, exactly two decimal places |
:.1% | Show as a percent (replace f with %), e.g. 0.065 → 6.5% |
:<9, :^9, :>9 | Left, centered, right aligned in a width of 9 characters |
- The format string starts with a colon and sits inside the closing brace, right after the variable or value.
- The
$is in the literal part, outside the braces, so alignment does not move it. - Multiline output: use
\nin the literal part (not inside braces), or put the f-string in triple quotes and break lines where you want them.
user1 = "Alberto"; user2 = "Babs"; user3 = "Carlos"
print(f"{user1} \n{user2} \n{user3}")Alberto Babs Carlos
Binary, octal and hexadecimal
| Base | Prefix | Convert 255 |
|---|---|---|
| Binary (base 2) | 0b | bin(255) → 0b11111111 |
| Octal (base 8) | 0o | oct(255) → 0o377 |
| Hexadecimal (base 16) | 0x | hex(255) → 0xff |
To convert to decimal you need no function. Just print the number: print(0b11111111), print(0o377) and print(0xff) all show 255.
Quick check
Difference between int(3.9) and round(3.9)?
int truncates the decimals (gives 3). round rounds to the nearest value (gives 4).
How do you use sqrt?
Import the math module first, then call math.sqrt(x).
What does f"{x:,.2f}" do?
Shows x with commas in thousands and exactly two decimal places.
What does bin(255) return?
The string 0b11111111.
Working with Text and Dates (Part 2)
Big picture. Strings can be joined, measured, sliced and changed with operators and methods. Dates and times need the datetime module, because Python has no built-in date type.
Manipulating strings
Concatenation and length
Join strings with + (concatenation). Python does not add spaces for you, so add " " yourself. len() counts characters, including spaces.
full_name = "Alan" + " " + "C" + ". " + "Simpson"
print(full_name)
print(len(""), len(" "), len("A B C"))Alan C. Simpson 0 1 5
String operators
| Operator | Purpose |
|---|---|
x in s / x not in s | True if x is / is not somewhere in s |
s * n | Repeat string s n times |
s[i] | Character at position i (the first character is 0) |
s[i:j] | Slice from position i up to position j |
s[i:j:k] | Slice from i to j with step k |
min(s) / max(s) | Smallest / largest character |
s.index(x) | Position of the first occurrence of x |
s.count(x) | How many times x appears in s |
Characters are compared by their ASCII numbers: space is 32, A to Z are 65 to 90, and a to z are 97 to 122. That is why max("Abc") is c.
String methods
| Method | Purpose |
|---|---|
s.upper() / s.lower() | All uppercase / all lowercase |
s.capitalize() | First letter capital, rest lowercase |
s.title() | First letter of every word capitalized |
s.swapcase() | Swap upper and lower case |
s.strip(), lstrip(), rstrip() | Remove spaces: both ends / leading / trailing |
s.replace(x, y) | Copy of s with every x replaced by y |
s.find(x) / s.rfind(x) | First position of x (forwards / backwards). Returns −1 if not found |
s.index(x) / s.rindex(x) | Like find, but gives an error if not found |
s.count(x) | Number of times x appears |
isalpha() | True if only letters |
isdecimal() / isnumeric() | True if only digits 0–9 |
islower(), isupper(), istitle(), isprintable() | True/False checks on the case or content |
find vs index.Both return the position of a substring. If it is missing, find gives −1 but index raises an error.
Dates and times
There is no built-in date type. Import the module (an alias is common):
import datetime as dtdatetime.date
Month, day and year. No time.
datetime.time
Hour, minute, second, microsecond (and optional time zone). No date.
datetime.datetime
Date and time together, plus optional time zone.
Working with dates
today = dt.date.today() # from the computer's clock
last_of_teens = dt.date(2019, 12, 31) # year, month, day, in that order
print(last_of_teens.month, last_of_teens.day, last_of_teens.year)12 31 2019
No leading zeros.Write April 1, 2020 as dt.date(2020, 4, 1). 2020, 04, 01 does not work.
Working with times and datetimes
right_now = dt.datetime.now()
print(right_now)2019-11-19 14:03:07.525975
Formatting dates with strftime directives
| Directive | Meaning | Directive | Meaning |
|---|---|---|---|
%Y / %y | 4-digit / 2-digit year | %j | Day number of the year (001–366) |
%m | Month number | %U / %W | Week number (Sunday / Monday first) |
%B / %b | Month name / abbreviated | %H / %I | Hour (24-hour / 12-hour) |
%d | Day of month | %M / %S | Minute / second |
%A / %a | Weekday name / abbreviated | %p | AM or PM |
%f | Microsecond | %z / %Z | UTC offset / time zone name |
%c / %x / %X | Local date-and-time / date / time | %% | A literal % character |
Example: the format %A %B %d is day number %j of %Y gives Saturday June 01 is day number 152 of 2019.
Calculating timespans with timedelta
A timedelta object is created automatically when you subtract two dates, times or datetimes. It measures “how long”, not “when”.
today = dt.date.today()
birthdate = dt.date(2000, 1, 31)
delta_age = today - birthdate # a timedelta
days_old = delta_age.days
years = days_old // 365 # floor division gives whole years
months = (days_old % 365) // 30 # leftover days, about 30 days per month
print(f"You are {years} years and {months} months old.")You are 18 years and 9 months old.
Time zones
Naive datetime
Has no time zone information.
Aware datetime
Includes time zone information.
The time from your system clock (now()) is in your own zone but does not say which. Compare it with UTC (utcnow()) and subtract to find your offset. To get the time in specific zones, use gettz from dateutil.tz:
import datetime as dt
from dateutil.tz import gettz
utc = dt.datetime.now(gettz('Etc/UTC'))
est = dt.datetime.now(gettz('America/New_York'))
print(f"{est:%A %D %I:%M %p %Z}")Quick check
What does len("A B C") return?
5, because spaces count as characters.
What is the index of the first character of a string?
0.
Which module handles dates, and what are its three main classes?
datetime: date, time and datetime.
What do you get when you subtract two dates?
A timedelta object (use .days for the number of days).
Naive vs aware datetime?
Naive has no time zone information; aware includes it.
Controlling the Action: Decisions with if
Big picture. Programs make decisions by comparing values. The if statement runs a block of code only when a condition is true, and else and elif handle the other cases.
Operators for making decisions
- Relational (comparison) operators
< <= > >= == !=compare two items to see how they are related. - Logical (Boolean) operators
and,or,notcombine several comparisons before a final decision.
Two equal signs.Test equality with ==, with no space between the signs. A single = is assignment.
The if statement
The condition is followed by a colon. Indent the lines that belong to the if. They run only if the condition is true. The first un-indented line after the block runs no matter what.
total = 100
sales_tax_rate = 0.065
taxable = True
if taxable:
print(f"Subtotal : ${total:.2f}")
sales_tax = total * sales_tax_rate
print(f"Sales Tax: ${sales_tax:.2f}")
total = total + sales_tax
print(f"Total : ${total:.2f}")Subtotal : $100.00 Sales Tax: $6.50 Total : $106.50
With taxable = False, all four indented lines are skipped and the total stays $100.00.
Adding else
Lines indented under else: run only if the condition was not true.
import datetime as dt
now = dt.datetime.now()
if now.hour < 12:
print("Good morning")
else:
print("Good afternoon")
print("I hope you are doing well!")Good morning I hope you are doing well!
But at 11 at night, “Good afternoon” is wrong. That is where elif comes in.
Handling many cases with elif
An if can have any number of elif conditions. A final else is optional and runs only if the if and all the previous elifs were false.
light_color = "yellow"
if light_color == "green":
print("Go")
elif light_color == "red":
print("Stop")
else:
print("Proceed with caution")
print("This code executes no matter what")Proceed with caution This code executes no matter what
Note: the contents list for Module 12 also mentions repeating a process with for and looping with while, but the slides in this file cover only Part 1 (decisions with if).
Quick check
What happens to the indented lines if the condition is false?
They are skipped. The un-indented code after the block still runs.
Difference between = and ==?
= assigns a value to a variable. == tests whether two values are equal.
When does else run?
Only when the if and all previous elif conditions are false.
Predict the output: light_color = "red" in the traffic-light code.
Stop and then This code executes no matter what.
Summary of the course slides for DS 231 · Based on Pierson, Data Science (3rd ed.) and Shovic & Simpson, Python All-in-One (2nd ed.)