import numpy as np8 NumPy 1D-Arrays
(CSE331) Python for Data Science
NumPy (Numerical Python) provides arrays and tools for numerical computing. Its array operations let us calculate with many values at once, usually without writing Python loops.
NumPy is a third-party package and is commonly included with Anaconda. Use the installation instructions from Lecture 7 if it is unavailable. By convention, import it as np:
1 Creating a One-Dimensional Array
The main NumPy object is the n-dimensional array, or ndarray. A 1D array stores a sequence of values; a 2D array arranges values in rows and columns.
my_list = [4, 1, 7, 3, 5]
my_array = np.array(my_list)
print(my_list)
print(my_array)
print(type(my_array))[4, 1, 7, 3, 5]
[4 1 7 3 5]
<class 'numpy.ndarray'>
For large numerical collections, NumPy arrays often use less memory and support faster calculations than Python lists. Each array has one data type, called its dtype.
1.1 Inspecting an array
print(my_array.ndim) # 1: number of dimensions
print(my_array.shape) # (5,): length of each dimension
print(my_array.size) # 5: total number of elements
print(my_array.dtype) # integer dtype; exact width can depend on the system1
(5,)
5
int64
These are attributes, so no parentheses are needed. The comma in (5,) identifies a one-item tuple. For a 1D array, len(my_array) and my_array.size are equal.
2 Data Types and Conversion
NumPy chooses a common dtype from the supplied values, or we can specify one.
integers = np.array([8, 4, 5, 2])
decimals = np.array([8, 4, 5, 2], dtype=float)
mixed_numbers = np.array([1, 2.5, 3])
print(integers.dtype, decimals.dtype)
print(mixed_numbers) # integers converted to floating-point valuesint64 float64
[1. 2.5 3. ]
Assigning a new value does not automatically change the array’s dtype:
integers[2] = 7.9
print(integers) # [8 4 7 2]: 7.9 was truncated to 7[8 4 7 2]
Use .astype() to convert the array before storing fractional values:
converted = integers.astype(float)
converted[2] = 7.9
print(converted)
print(integers) # original array remains integer-valued[8. 4. 7.9 2. ]
[8 4 7 2]
Converting afterward cannot recover a fractional part that has already been discarded. A value that cannot be converted, such as "hello" assigned to an integer array, raises an error.
3 Indexing and Slicing
Indices start at 0. A negative index counts from the end. A slice excludes its stop position.
x = np.array([10, 20, 30, 40, 50, 60])
print(x[0], x[-1])
print(x[1:4]) # [20 30 40]
print(x[::2]) # [10 30 50]
print(x[::-1])10 60
[20 30 40]
[10 30 50]
[60 50 40 30 20 10]
Assign to a position or slice to change existing values:
x[0] = 15
x[1:3] = [25, 35]
print(x)[15 25 35 40 50 60]
3.1 Views and copies
Unlike a Python list slice, a basic NumPy slice shares data with the original array. Such an array is called a view.
a = np.array([10, 20, 30, 40])
part = a[1:3]
part[0] = 99
print(a) # [10 99 30 40][10 99 30 40]
Use .copy() when changes should be independent:
separate = a[1:3].copy()
separate[0] = -1
print(separate) # [-1 30]
print(a) # unchanged by the preceding assignment[-1 30]
[10 99 30 40]
b = a gives another name for the same array; b = a.copy() creates independent numerical data.
4 Elementwise Arithmetic
A list comprehension applies a calculation to each item:
values = [4, 1, 7, 3, 5]
print([5 * value for value in values])[20, 5, 35, 15, 25]
An array expresses the same calculation directly:
a = np.array(values)
print(5 * a)
print(a + 100)
print(a**2)[20 5 35 15 25]
[104 101 107 103 105]
[16 1 49 9 25]
This is vectorized arithmetic: NumPy applies the operation to the elements. Recall that 5 * values repeats a Python list.
4.1 Two arrays
Arrays with the same shape combine element by element:
a = np.array([1, 4, 3])
b = np.array([5, 8, 2])
print(a + b)
print(a - b)
print(a * b)
print(a / b)[ 6 12 5]
[-4 -4 1]
[ 5 32 6]
[0.2 0.5 1.5]
For two 1D arrays, their lengths must match, or one must have length 1. A scalar also works. These are simple cases of broadcasting, covered in Lecture 9.
print(a + np.array([10])) # [11 14 13][11 14 13]
Lengths 3 and 4 are incompatible:
np.array([2, 1, 4]) + np.array([3, 9, 2, 7])--------------------------------------------------------------------------- ValueError Traceback (most recent call last) Cell In[15], line 1 ----> 1 np.array([2, 1, 4]) + np.array([3, 9, 2, 7]) ValueError: operands could not be broadcast together with shapes (3,) (4,)
This example deliberately produces a ValueError.
5 Creating Special Arrays
print(np.zeros(5))
print(np.ones(5))
print(np.zeros(5, dtype=int))[0. 0. 0. 0. 0.]
[1. 1. 1. 1. 1.]
[0 0 0 0 0]
zeros() and ones() use a floating-point dtype by default.
| Function | Specify | Endpoint behaviour |
|---|---|---|
np.arange(start, stop, step) |
Step size | Intended to exclude stop |
np.linspace(start, stop, num) |
Number of values | Includes stop by default |
print(np.arange(2, 10, 2))
print(np.arange(2, 4, 0.25))
print(np.linspace(2, 4, 11))[2 4 6 8]
[2. 2.25 2.5 2.75 3. 3.25 3.5 3.75]
[2. 2.2 2.4 2.6 2.8 3. 3.2 3.4 3.6 3.8 4. ]
Floating-point steps in arange() can produce rounding surprises. Use linspace() when the required number of points and endpoints are important. Set endpoint=False if the final endpoint should be excluded.
6 Array Functions
6.1 Summaries
x = np.array([3.2, 4.8, 8.7, 8.7, 6.4, 5.3])
print("Sum:", np.sum(x))
print("Product:", np.prod(x))
print("Minimum and maximum:", np.min(x), np.max(x))
print("Mean and median:", np.mean(x), np.median(x))
print("Index of minimum:", np.argmin(x))
print("Index of maximum:", np.argmax(x))
print("Distinct values:", np.unique(x))Sum: 37.099999999999994
Product: 39435.33772799999
Minimum and maximum: 3.2 8.7
Mean and median: 6.183333333333333 5.85
Index of minimum: 0
Index of maximum: 2
Distinct values: [3.2 4.8 5.3 6.4 8.7]
argmin() and argmax() give positions, not values. In a 1D array, a tie returns the position of the first occurrence. unique() returns sorted distinct values for these numerical arrays.
Many summaries are also array methods: x.mean() and np.mean(x) give the same result here.
6.2 Standard deviation: check the denominator
NumPy’s default standard deviation uses denominator \(n\). For the usual sample standard deviation, use ddof=1 to obtain denominator \(n-1\):
\[ s = \sqrt{\frac{\sum_{i=1}^{n}(x_i-\bar{x})^2}{n-1}}. \]
print("SD using n:", np.std(x))
print("Sample SD:", np.std(x, ddof=1))SD using n: 2.0128062223892513
Sample SD: 2.2049187437787054
R’s sd() uses the sample convention. At least two observations are required for ddof=1.
6.3 Elementwise functions
These functions transform individual values:
positive = np.array([1.0, 2.0, 4.0])
print(np.sqrt(positive))
print(np.exp(positive))
print(np.log(positive)) # natural logarithms
print(np.round(np.log(positive), 2))[1. 1.41421356 2. ]
[ 2.71828183 7.3890561 54.59815003]
[0. 0.69314718 1.38629436]
[0. 0.69 1.39]
For finite real logarithms, use positive inputs. np.log(0) gives -inf and negative real inputs give nan, with warnings. Round for display when appropriate; retain full precision in intermediate calculations.
7 Comparisons and Boolean Masks
A comparison returns one Boolean value for each element:
values = np.array([4, 7, 6, 3, 9, 8])
mask = values > 5
print(mask)
print(values[mask]) # [7 6 9 8][False True True False True True]
[7 6 9 8]
Selecting values with a Boolean array is Boolean masking. For this 1D selection, the mask must have the same length as the array.
print(np.sum(mask)) # 4: count of True values
print(np.mean(mask)) # 4/6: proportion meeting the condition
print(np.any(mask)) # True if at least one value meets the condition
print(np.all(mask)) # True if every value meets the condition4
0.6666666666666666
True
False
The proportion interpretation assumes a nonempty array. A Boolean counts as 1 for True and 0 for False in these calculations.
7.1 Combining conditions
Use &, |, and ~ on Boolean arrays. Put each comparison in parentheses.
| Operation | Example |
|---|---|
| AND | (values > 5) & (values < 9) |
| OR | (values < 4) | (values > 8) |
| NOT | ~(values > 5) |
print(values[(values > 5) & (values < 9)])
print(values[(values < 4) | (values > 8)])
print(values[~(values > 5)])[7 6 8]
[3 9]
[4 3]
Do not combine array comparisons with Python’s and or or, or write 5 < values < 9. A multi-element Boolean array also cannot be used directly as an if condition; reduce it with np.any() or np.all() when one Boolean decision is needed.
7.2 Selecting and updating
scores = np.array([72, -1, 0, 84, -1])
selected = scores[scores >= 0]
print(selected)
scores[scores < 0] = 0
print(scores)[72 0 84]
[72 0 0 84 0]
Boolean selection creates a copy. Direct indexed assignment, as in the last operation, changes the original array. Here replacing negatives by zero is only an assignment demonstration; a missing score should not normally be treated as a recorded zero.
8 Selecting by Position: Fancy Indexing
Supply a list or array of integer positions to select items in a chosen order:
x = np.array([10, 20, 30, 40, 50, 60, 70, 80, 90])
print(x[[6, 3, 8]]) # [70 40 90][70 40 90]
Integer-array selection creates a copy, just like Boolean selection. Basic slicing, such as x[1:4], creates a view.
9 Worked Example: Comparing Groups
The category and value arrays describe the same observations in the same order.
category = np.array(["A", "B", "A", "C", "B", "A", "C", "A"])
value = np.array([8, 6, 3, 1, 10, 12, 6, 4])
group_a = value[category == "A"]
print("Group A values:", group_a)
print("Group A count:", group_a.size)
print("Group A mean:", np.mean(group_a))
print("Group A values above 5:", value[(category == "A") & (value > 5)])Group A values: [ 8 3 12 4]
Group A count: 4
Group A mean: 6.75
Group A values above 5: [ 8 12]
The mean is 6.75. A combined mask can select values or be summed to count observations. Check that a selected group is nonempty before calculating its mean.
10 Missing Numerical Values: A First Look
Use np.nan to represent missing values in a floating-point array. It may also appear after an undefined numerical calculation.
recorded = np.array([72, np.nan, 0, 84])
print(np.isnan(recorded))
print(np.mean(recorded)) # nan
print(np.nanmean(recorded)) # 52.0: excludes nan, retains zero[False True False False]
nan
52.0
Use np.isnan(), not == np.nan, to identify these entries. An integer array cannot store np.nan; create or convert to a floating-point array first. If all entries are nan, nanmean() returns nan with a warning.
11 Exercises
- Create
x = np.array([12, 7, 18, 5, 9, 14]). Display its shape, size, dtype, first element, last element, and first three elements. - Create separate arrays containing
x + 5,x**2, andx / 2. - Create the even integers from 2 through 20 using
arange(), then create 9 evenly spaced values from 0 to 1 inclusive usinglinspace(). - Calculate the mean, median, and sample standard deviation of
x. Find the index of its maximum. - Select values in
xthat are at least 8 and below 15. Count them and find their proportion. - Select the last, first, and third elements of
xin that order using integer indices. - Create a slice of
xand modify it. Repeat using.copy()and explain the difference. - For the worked group data, find the mean of group B and count observations in group A or C whose value is at least 6.
- For
y = np.array([10, np.nan, 14, 0, np.nan]), count missing entries and compute the mean of recorded values.
12 Summary
- Array arithmetic is elementwise; inspect
shapeanddtypebefore calculations. - Basic slices share data; use
.copy()for independent data. - Use
ddof=1for the usual sample standard deviation. - Boolean masks select observations;
&,|, and~combine conditions. - Distinguish missing values from zero.