li1 = [47, 52, 71, 90]
li1, type(li1), len(li1)([47, 52, 71, 90], list, 4)
(CSE331) Python for Data Science
x = [1, 2, 3]t = (1, 2, 3)r = range(1, 11)d = {"name": "Rasel", "dept": "IASDS"}s = {1, 2, 3}A list stores an ordered sequence of items. Create with square brackets:
li1 = [47, 52, 71, 90]
li1, type(li1), len(li1)([47, 52, 71, 90], list, 4)
li2 = [1, "a", [3, 4]]
li2, len(li2)([1, 'a', [3, 4]], 3)
47 in li1, 99 in li1(True, False)
List indices start at 0. An index selects one item; a slice selects a range of items.
marks = [58, 72, 81, 65]
print(marks[0], marks[-1])
marks[1] = 75
print(marks)58 65
[58, 75, 81, 65]
marks[4] would raise IndexError: the valid nonnegative indices here are 0 to 3.
print(sum(marks), min(marks), max(marks))
print(sum(marks) / len(marks))279 58 81
69.75
The mean calculation requires a nonempty collection.
int, float, str, list).object.method(...).append, extend, insert, pop, sort, clear, copy.li1.append(24) # add a single item to the end
li1[47, 52, 71, 90, 24]
li1.insert(2, 100) # insert at index 2
li1[47, 52, 100, 71, 90, 24]
li1.pop(1) # remove-and-return item at index 1
li1[47, 100, 71, 90, 24]
append(x) adds one element; extend(iterable) adds many elements:x = [4, None, "foo"]
x.extend([7, 8, (2, 3)]) # concatenates elements from the iterable
x[4, None, 'foo', 7, 8, (2, 3)]
[4, None, "foo"] + [7, 8, (2, 3)][4, None, 'foo', 7, 8, (2, 3)]
remove(value) removes the first matching value; pop(index) removes an item by position and returns it.
values = [5, 8, 5, 12]
values.remove(5)
removed = values.pop()
print(values, removed)[8, 5] 12
Methods such as append(), extend(), and sort() modify a list and return None. Do not write values = values.append(3).
reverse= and key=:a = [1, 2, 3, 31, 32, 33, 11, 12, 13]
a.sort()
a[1, 2, 3, 11, 12, 13, 31, 32, 33]
a.sort(reverse=True)
a[33, 32, 31, 13, 12, 11, 3, 2, 1]
names = ["alice", "Bob", "carol"]
names.sort(key=str.lower) # case-insensitive sort
names['alice', 'Bob', 'carol']
b = [3, 1, 2]
sorted_b = sorted(b, reverse=True)
b, sorted_b([3, 1, 2], [3, 2, 1])
seq[start:stop] returns a new list; start inclusive, stop exclusive.seq = [1, 2, 3, 41, 42, 43, 11, 12, 13]
seq[3:4], seq[3:5]([41], [41, 42])
start or stop to use defaults:seq[:5], seq[5:]([1, 2, 3, 41, 42], [43, 11, 12, 13])
seq[-4:], seq[-6:-2]([43, 11, 12, 13], [41, 42, 43, 11])
seq[3:5] = [6, 3]
seq[1, 2, 3, 6, 3, 43, 11, 12, 13]
seq[start:stop:step] (can be negative):seq[::2], seq[::-1] # every 2nd; reversed copy([1, 3, 3, 11, 13], [13, 12, 11, 43, 3, 6, 3, 2, 1])
R.a = 1947
b = a # both refer to the same int object (1947)
a = 1971 # a now refers to a new int object; b still 1947
print(a, b)1971 1947
orig = [1, 2, 3]
alias = orig # same object
orig.append(99)
alias, orig([1, 2, 3, 99], [1, 2, 3, 99])
int, float, bool, str) are immutable; when you “change” them, you are actually re-binding the name to a new object.list are often mutable; mutation changes the object itself without re-binding the name.orig = [1, 2, 3]
alias = orig # same object
shallow = orig[:] # new list (shallow copy): or: orig.copy(), list(orig)
orig.append(99)
alias, shallow([1, 2, 3, 99], [1, 2, 3])
a = [1, 2]
b = [1, 2]
c = a
print(a == b) # True: equal values
print(a is b) # False: different lists
print(a is c) # True: same listTrue
False
True
rows = [[1, 2], [3, 4]]
print(rows[1][0]) # second row, first item
copied_rows = rows.copy()
copied_rows[0][0] = 99
print(rows) # inner list is shared3
[[99, 2], [3, 4]]
A shallow copy creates a new outer collection but shares its contained objects. This matters for nested data. For a flat list of immutable numbers, a shallow copy is normally enough.
Python list arithmetic is not elementwise:
print([1, 2] + [3, 4]) # concatenation
print([1, 2] * 2) # repetition[1, 2, 3, 4]
[1, 2, 1, 2]
NumPy arrays will provide elementwise numerical operations later.
A tuple is an immutable, ordered sequence.
tup1 = (4, 5, 6)
tup2 = 4, 5, 6 # parentheses optional in many contexts
tup1, tup2((4, 5, 6), (4, 5, 6))
A one-item tuple needs a trailing comma: (5,). (5) is simply the integer 5 in parentheses.
nested_tup = (4, 5, 6), (7, 8)
nested_tup((4, 5, 6), (7, 8))
tuple([4, 0, 2]), tuple("string")((4, 0, 2), ('s', 't', 'r', 'i', 'n', 'g'))
tup = ("pre", "post", 3)
tup[0], tup[-1]('pre', 3)
(4, None, "foo") + (6, 0) + ("bar",), ("pre", "post") * 3((4, None, 'foo', 6, 0, 'bar'), ('pre', 'post', 'pre', 'post', 'pre', 'post'))
Tuples are immutable, but they can contain mutable objects:
t = (1, [2, 3])
t[1].append(4) # allowed: we mutate the list inside the tuple
t(1, [2, 3, 4])
What’s immutable is the container (which objects are stored), not necessarily the contents.
x, y = (10, 20)
x, y(10, 20)
a, *b, c = (1, 2, 3, 4, 5)
a, b, c(1, [2, 3, 4], 5)
range is an immutable sequence of evenly spaced integers.range(start, stop, step)
start → first value (default = 0)stop → sequence goes up to but not including this valuestep → increment (default = 1)start, stop, and step, and computes elements on demand.len(), indexing, slicing, and membership testing.r = range(10) # 0...9
len(r), r[0], r[-1](10, 0, 9)
list(range(0, 20, 2))[0, 2, 4, 6, 8, 10, 12, 14, 16, 18]
list(range(10, 0, -3))[10, 7, 4, 1]
5 in range(10), 25 in range(10) # fast membership check(True, False)
range(0, 10, 2)[2:5] # slicing a range yields another rangerange(4, 10, 2)
dict is likely the most important built-in Python data structure. A more common name for it is hash map or associative array. It is a flexibly sized collection of key-value pairs, where key and value are Python objects.
Keys must be hashable: strings and numbers are common choices; lists cannot be keys.
Keys are unique. Assigning an existing key replaces its value.
Dictionaries preserve insertion order, but that does not mean keys are sorted.
One approach for creating a dict is to use curly braces {} and colons to separate keys and values:
empty_dict = {}
d1 = {'a' : 'some value', 'b' : [1, 2, 3, 4]}
d1{'a': 'some value', 'b': [1, 2, 3, 4]}
d1["b"][1, 2, 3, 4]
d1["dept"] = "IASDS"
d1["b"] = list(range(3))
d1{'a': 'some value', 'b': [0, 1, 2], 'dept': 'IASDS'}
"a" in d1, "z" in d1(True, False)
d1["missing"] would throw an error (KeyError) if the key missing doesn’t exist. d1.get("missing") avoids that error and gives you a safe fallback (None or the value you provide).d1.get("b") # [0, 1, 2] (key exists)
d1.get("age") # None (key missing)
d1.get("age", "N/A") # N/A (key missing, fallback given) 'N/A'
d1["dummy"] = "temp"
val = d1.pop("dummy") # returns value and removes key
val, d1('temp', {'a': 'some value', 'b': [0, 1, 2], 'dept': 'IASDS'})
d1["x"] = 1
del d1["x"] # delete by key (raises KeyError if absent)
d1{'a': 'some value', 'b': [0, 1, 2], 'dept': 'IASDS'}
keys(), values(), items():list(d1.keys()), list(d1.values()), list(d1.items())(['a', 'b', 'dept'],
['some value', [0, 1, 2], 'IASDS'],
[('a', 'some value'), ('b', [0, 1, 2]), ('dept', 'IASDS')])
d2 = {"b": 999, "new": 1}
d1.update(d2) # overwrites existing keys
d1{'a': 'some value', 'b': 999, 'dept': 'IASDS', 'new': 1}
A dictionary can describe one observation. A list of dictionaries can hold several observations.
students = [
{"name": "Ayesha", "score": 72},
{"name": "Imran", "score": None},
{"name": "Nadia", "score": 0},
]
print(students[0]["name"])
print(students[1]["score"] is None)Ayesha
True
Here None means an unrecorded score; 0 is a recorded score. We will process these records using loops, then return to tabular data in pandas.
A set stores unique, hashable elements without positional indexing. Great for membership tests and set algebra.
Use set() for an empty set; {} creates an empty dictionary. Do not depend on the display or iteration order of a set. Use sorted(s) when you need a reproducible sorted list of comparable values.
set([2, 2, 2, 1, 3, 3]), {2, 2, 2, 1, 3, 3} # literal deduplicates too({1, 2, 3}, {1, 2, 3})
a = {1, 2, 3, 4, 5}
b = {3, 4, 5, 6, 7, 8}
a | b, a & b, a - b, a ^ b # union, intersection, difference, symmetric diff({1, 2, 3, 4, 5, 6, 7, 8}, {3, 4, 5}, {1, 2}, {1, 2, 6, 7, 8})
s = {1, 2, 3}
s.add(4) # add single element
s.update([3, 5]) # add many
s{1, 2, 3, 4, 5}
s.remove(3) # remove; raises KeyError if absent
s.discard(3) # remove; does nothing if absent
s{1, 2, 4, 5}
| Function | Operator | Description |
|---|---|---|
a.add(x) |
Not applicable | Add element x to a |
a.clear() |
Not applicable | Remove all elements |
a.remove(x) |
Not applicable | Remove x (error if absent) |
a.discard(x) |
Not applicable | Remove x (no error if absent) |
a.pop() |
Not applicable | Remove & return an arbitrary element |
a.union(b) |
a \| b |
All elements in a or b |
a.update(b) |
a \|= b |
In-place union |
a.intersection(b) |
a & b |
Elements in both |
a.intersection_update(b) |
a &= b |
In-place intersection |
a.difference(b) |
a - b |
Elements in a not in b |
a.difference_update(b) |
a -= b |
In-place difference |
a.symmetric_difference(b) |
a ^ b |
In a or b but not both |
a.symmetric_difference_update(b) |
a ^= b |
In-place symmetric diff |
a.issubset(b) |
a <= b |
All elems of a in b |
a.issuperset(b) |
a >= b |
All elems of b in a |
a.isdisjoint(b) |
Not applicable | No elements in common |
In Python, different data types (list, tuple, range, set, dict) support different sets of methods. To find out which ones are available and how to use them:
dir(type) → lists all attributes (including methods) of that type.dir(list) # attributes and methods of list
dir(tuple) # attributes and methods of tupleappend, remove, add, keys).__len__, __add__) are special methods, usually ignored in day-to-day use.help(type.method) → shows detailed documentation of a specific method.help(list.append) # details about append()
help(dict.update) # details about update()x = [9, 2, 7, 2, 5], obtain the first two items, the last two items, and the reversed list.x without modifying x. Then append 10 to the original.b = a differs from b = a.copy() for a flat list of numbers. Demonstrate with an item assignment.(12, 15) into two variables.name and score; update the score and retrieve a missing email with the fallback "Unknown".group_a = {1, 2, 4, 5} and group_b = {2, 3, 5}, find their common members and members present only in group A.[1, 2, 3] * 2 does not double each value.append/extend, remove/discard.