Commits · 3ff81ad2dedc7cc3defb6e418d7c5fb415d56026 · cs525-sp18-g07 / spark

Aug 17, 2015

[SPARK-9199] [CORE] Upgrade Tachyon version from 0.7.0 -> 0.7.1. · 3ff81ad2

Calvin Jia authored 9 years ago

Updates the tachyon-client version to the latest release.

The main difference between 0.7.0 and 0.7.1 on the client side is to support running Tachyon on local file system by default.

No new non-Tachyon dependencies are added, and no code changes are required since the client API has not changed.

Author: Calvin Jia <jia.calvin@gmail.com>

Closes #8235 from calvinjia/spark-9199-master.

3ff81ad2

[SPARK-9871] [SPARKR] Add expression functions into SparkR which have a variable parameter · 26e76058

Yu ISHIKAWA authored 9 years ago

### Summary

- Add `lit` function
- Add `concat`, `greatest`, `least` functions

I think we need to improve `collect` function in order to implement `struct` function. Since `collect` doesn't work with arguments which includes a nested `list` variable. It seems that a list against `struct` still has `jobj` classes. So it would be better to solve this problem on another issue.

### JIRA
[[SPARK-9871] Add expression functions into SparkR which have a variable parameter - ASF JIRA](https://issues.apache.org/jira/browse/SPARK-9871)

Author: Yu ISHIKAWA <yuu.ishikawa@gmail.com>

Closes #8194 from yu-iskw/SPARK-9856.

26e76058

Aug 16, 2015

[SPARK-10005] [SQL] Fixes schema merging for nested structs · ae2370e7

Cheng Lian authored 9 years ago

In case of schema merging, we only handled first level fields when converting Parquet groups to `InternalRow`s. Nested struct fields are not properly handled.

For example, the schema of a Parquet file to be read can be:

```
message individual {
  required group f1 {
    optional binary f11 (utf8);
  }
}
```

while the global schema is:

```
message global {
  required group f1 {
    optional binary f11 (utf8);
    optional int32 f12;
  }
}
```

This PR fixes this issue by padding missing fields when creating actual converters.

Author: Cheng Lian <lian@databricks.com>

Closes #8228 from liancheng/spark-10005/nested-schema-merging.

ae2370e7

[SPARK-10008] Ensure shuffle locality doesn't take precedence over narrow deps · cf016075

Matei Zaharia authored 9 years ago

The shuffle locality patch made the DAGScheduler aware of shuffle data,
but for RDDs that have both narrow and shuffle dependencies, it can
cause them to place tasks based on the shuffle dependency instead of the
narrow one. This case is common in iterative join-based algorithms like
PageRank and ALS, where one RDD is hash-partitioned and one isn't.

Author: Matei Zaharia <matei@databricks.com>

Closes #8220 from mateiz/shuffle-loc-fix.

cf016075

[SPARK-8844] [SPARKR] head/collect is broken in SparkR. · 5f9ce738

Sun Rui authored 9 years ago

This is a WIP patch for SPARK-8844  for collecting reviews.

This bug is about reading an empty DataFrame. in readCol(),
      lapply(1:numRows, function(x) {
does not take into consideration the case where numRows = 0.

Will add unit test case.

Author: Sun Rui <rui.sun@intel.com>

Closes #7419 from sun-rui/SPARK-8844.

5f9ce738

[SPARK-9973] [SQL] Correct in-memory columnar buffer size · 182f9b7a

Kun Xu authored 9 years ago

The `initialSize` argument of `ColumnBuilder.initialize()` should be the
number of rows rather than bytes.  However `InMemoryColumnarTableScan`
passes in a byte size, which makes Spark SQL allocate more memory than
necessary when building in-memory columnar buffers.

Author: Kun Xu <viper_kun@163.com>

Closes #8189 from viper-kun/errorSize.

182f9b7a

Aug 15, 2015

[SPARK-9805] [MLLIB] [PYTHON] [STREAMING] Added _eventually for ml streaming pyspark tests · 1db7179f

Joseph K. Bradley authored 9 years ago

Recently, PySpark ML streaming tests have been flaky, most likely because of the batches not being processed in time. Proposal: Replace the use of _ssc_wait (which waits for a fixed amount of time) with a method which waits for a fixed amount of time but can terminate early based on a termination condition method. With this, we can extend the waiting period (to make tests less flaky) but also stop early when possible (making tests faster on average, which I verified locally).

CC: mengxr tdas freeman-lab

Author: Joseph K. Bradley <joseph@databricks.com>

Closes #8087 from jkbradley/streaming-ml-tests.

1db7179f

[SPARK-9955] [SQL] correct error message for aggregate · 57056725

Wenchen Fan authored 9 years ago

We should skip unresolved `LogicalPlan`s for `PullOutNondeterministic`, as calling `output` on unresolved `LogicalPlan` will produce confusing error message.

Author: Wenchen Fan <cloud0fan@outlook.com>

Closes #8203 from cloud-fan/error-msg and squashes the following commits:

1c67ca7 [Wenchen Fan] move test
7593080 [Wenchen Fan] correct error message for aggregate

57056725

[SPARK-9980] [BUILD] Fix SBT publishLocal error due to invalid characters in doc · a85fb6c0

Herman van Hovell authored 9 years ago

Tiny modification to a few comments ```sbt publishLocal``` work again.

Author: Herman van Hovell <hvanhovell@questtec.nl>

Closes #8209 from hvanhovell/SPARK-9980.

a85fb6c0

[SPARK-9725] [SQL] fix serialization of UTF8String across different JVM · 7c1e5682

Davies Liu authored 9 years ago

The BYTE_ARRAY_OFFSET could be different in JVM with different configurations (for example, different heap size, 24 if heap > 32G, otherwise 16), so offset of UTF8String is not portable, we should handler that during serialization.

Author: Davies Liu <davies@databricks.com>

Closes #8210 from davies/serialize_utf8string.

7c1e5682

Aug 14, 2015

[SPARK-9960] [GRAPHX] sendMessage type fix in LabelPropagation.scala · 71a3af8a
zc he authored 9 years ago
```
Author: zc he <farseer90718@gmail.com>

Closes #8188 from farseer90718/farseer-patch-1.
```
71a3af8a

[SPARK-9984] [SQL] Create local physical operator interface. · 609ce3c0

Reynold Xin authored 9 years ago

This pull request creates a new operator interface that is more similar to traditional database query iterators (with open/close/next/get).

These local operators are not currently used anywhere, but will become the basis for SPARK-9983 (local physical operators for query execution).

cc zsxwing

Author: Reynold Xin <rxin@databricks.com>

Closes #8212 from rxin/SPARK-9984.

609ce3c0

[SPARK-8887] [SQL] Explicit define which data types can be used as dynamic partition columns · 6c4fdbec

Yijie Shen authored 9 years ago

This PR enforce dynamic partition column data type requirements by adding analysis rules.

JIRA: https://issues.apache.org/jira/browse/SPARK-8887

Author: Yijie Shen <henry.yijieshen@gmail.com>

Closes #8201 from yjshen/dynamic_partition_columns.

6c4fdbec

[SPARK-9634] [SPARK-9323] [SQL] cleanup unnecessary Aliases in LogicalPlan at the end of analysis · ec29f203

Wenchen Fan authored 9 years ago

Also alias the ExtractValue instead of wrapping it with UnresolvedAlias when resolve attribute in LogicalPlan, as this alias will be trimmed if it's unnecessary.

Based on #7957 without the changes to mllib, but instead maintaining earlier behavior when using `withColumn` on expressions that already have metadata.

Author: Wenchen Fan <cloud0fan@outlook.com>
Author: Michael Armbrust <michael@databricks.com>

Closes #8215 from marmbrus/pr/7957.

ec29f203

[HOTFIX] fix duplicated braces · 37586e54

Davies Liu authored 9 years ago

Author: Davies Liu <davies@databricks.com>

Closes #8219 from davies/fix_typo.

37586e54

[SPARK-9934] Deprecate NIO ConnectionManager. · e5fd6041

Reynold Xin authored 9 years ago

Deprecate NIO ConnectionManager in Spark 1.5.0, before removing it in Spark 1.6.0.

Author: Reynold Xin <rxin@databricks.com>

Closes #8162 from rxin/SPARK-9934.

e5fd6041

[SPARK-9949] [SQL] Fix TakeOrderedAndProject's output. · 932b24fd

Yin Huai authored 9 years ago

https://issues.apache.org/jira/browse/SPARK-9949

Author: Yin Huai <yhuai@databricks.com>

Closes #8179 from yhuai/SPARK-9949.

932b24fd

[SPARK-9968] [STREAMING] Reduced time spent within synchronized block to prevent lock starvation · 18a761ef

Tathagata Das authored 9 years ago

When the rate limiter is actually limiting the rate at which data is inserted into the buffer, the synchronized block of BlockGenerator.addData stays blocked for long time. This causes the thread switching the buffer and generating blocks (synchronized with addData) to starve and not generate blocks for seconds. The correct solution is to not block on the rate limiter within the synchronized block for adding data to the buffer.

Author: Tathagata Das <tathagata.das1565@gmail.com>

Closes #8204 from tdas/SPARK-9968 and squashes the following commits:

8cbcc1b [Tathagata Das] Removed unused val
a73b645 [Tathagata Das] Reduced time spent within synchronized block

18a761ef

[SPARK-9966] [STREAMING] Handle couple of corner cases in PIDRateEstimator · f3bfb711

Tathagata Das authored 9 years ago

1. The rate estimator should not estimate any rate when there are no records in the batch, as there is no data to estimate the rate. In the current state, it estimates and set the rate to zero. That is incorrect.

2. The rate estimator should not never set the rate to zero under any circumstances. Otherwise the system will stop receiving data, and stop generating useful estimates (see reason 1). So the fix is to define a parameters that sets a lower bound on the estimated rate, so that the system always receives some data.

Author: Tathagata Das <tathagata.das1565@gmail.com>

Closes #8199 from tdas/SPARK-9966 and squashes the following commits:

829f793 [Tathagata Das] Fixed unit test and added comments
3a994db [Tathagata Das] Added min rate and updated tests in PIDRateEstimator

f3bfb711

[SPARK-8670] [SQL] Nested columns can't be referenced in pyspark · 1150a19b

Wenchen Fan authored 9 years ago

This bug is caused by a wrong column-exist-check in `__getitem__` of pyspark dataframe. `DataFrame.apply` accepts not only top level column names, but also nested column name like `a.b`, so we should remove that check from `__getitem__`.

Author: Wenchen Fan <cloud0fan@outlook.com>

Closes #8202 from cloud-fan/nested.

1150a19b

[SPARK-9981] [ML] Made labels public for StringIndexerModel · 2a6590e5

Joseph K. Bradley authored 9 years ago

Also added unit test for integration between StringIndexerModel and IndexToString

CC: holdenk We realized we should have left in your unit test (to catch the issue with removing the inverse() method), so this adds it back. mengxr

Author: Joseph K. Bradley <joseph@databricks.com>

Closes #8211 from jkbradley/stridx-labels.

2a6590e5

[SPARK-9978] [PYSPARK] [SQL] fix Window.orderBy and doc of ntile() · 11ed2b18
Davies Liu authored 9 years ago
```
Author: Davies Liu <davies@databricks.com>

Closes #8213 from davies/fix_window.
```
11ed2b18

[SPARK-9877] [CORE] Fix StandaloneRestServer NPE when submitting application · 9407baa2

jerryshao authored 9 years ago

Detailed exception log can be seen in [SPARK-9877](https://issues.apache.org/jira/browse/SPARK-9877), the problem is when creating `StandaloneRestServer`, `self` (`masterEndpoint`) is null. So this fix is creating `StandaloneRestServer` when `self` is available.

Author: jerryshao <sshao@hortonworks.com>

Closes #8127 from jerryshao/SPARK-9877.

9407baa2

[SPARK-9948] Fix flaky AccumulatorSuite - internal accumulators · 6518ef63

Andrew Or authored 9 years ago

In these tests, we use a custom listener and we assert on fields in the stage / task completion events. However, these events are posted in a separate thread so they're not guaranteed to be posted in time. This commit fixes this flakiness through a job end registration callback.

Author: Andrew Or <andrew@databricks.com>

Closes #8176 from andrewor14/fix-accumulator-suite.

6518ef63

[SPARK-9809] Task crashes because the internal accumulators are not properly initialized · 33bae585

Carson Wang authored 9 years ago

When a stage failed and another stage was resubmitted with only part of partitions to compute, all the tasks failed with error message: java.util.NoSuchElementException: key not found: peakExecutionMemory.
This is because the internal accumulators are not properly initialized for this stage while other codes assume the internal accumulators always exist.

Author: Carson Wang <carson.wang@intel.com>

Closes #8090 from carsonwang/SPARK-9809.

33bae585

[SPARK-9828] [PYSPARK] Mutable values should not be default arguments · ffa05c84
MechCoder authored 9 years ago
```
Author: MechCoder <manojkumarsivaraj334@gmail.com>

Closes #8110 from MechCoder/spark-9828.
```
ffa05c84

[SPARK-9561] Re-enable BroadcastJoinSuite · ece00566

Andrew Or authored 9 years ago

We can do this now that SPARK-9580 is resolved.

Author: Andrew Or <andrew@databricks.com>

Closes #8208 from andrewor14/reenable-sql-tests.

ece00566

[SPARK-9946] [SPARK-9589] [SQL] fix NPE and thread-safety in TaskMemoryManager · 3bc55287

Davies Liu authored 9 years ago

Currently, we access the `page.pageNumer` after it's freed, that could be modified by other thread, cause NPE.

The same TaskMemoryManager could be used by multiple threads (for example, Python UDF and TransportScript), so it should be thread safe to allocate/free memory/page. The underlying Bitset and HashSet are not thread safe, we should put them inside a synchronized block.

cc JoshRosen

Author: Davies Liu <davies@databricks.com>

Closes #8177 from davies/memory_manager.

3bc55287

[SPARK-9923] [CORE] ShuffleMapStage.numAvailableOutputs should be an Int instead of Long · 57c2d088

Neelesh Srinivas Salian authored 9 years ago

Modified type of ShuffleMapStage.numAvailableOutputs from Long to Int

Author: Neelesh Srinivas Salian <nsalian@cloudera.com>

Closes #8183 from nssalian/SPARK-9923.

57c2d088

[SPARK-9929] [SQL] support metadata in withColumn · 34d610be

Wenchen Fan authored 9 years ago

in MLlib sometimes we need to set metadata for the new column, thus we will alias the new column with metadata before call `withColumn` and in `withColumn` we alias this clolumn again. Here I overloaded `withColumn` to allow user set metadata, just like what we did for `Column.as`.

Author: Wenchen Fan <cloud0fan@outlook.com>

Closes #8159 from cloud-fan/withColumn.

34d610be

[SPARK-8744] [ML] Add a public constructor to StringIndexer · a7317ccd

Holden Karau authored 9 years ago

It would be helpful to allow users to pass a pre-computed index to create an indexer, rather than always going through StringIndexer to create the model.

Author: Holden Karau <holden@pigscanfly.ca>

Closes #7267 from holdenk/SPARK-8744-StringIndexerModel-should-have-public-constructor.

a7317ccd

[SPARK-9956] [ML] Make trees work with one-category features · 7ecf0c46

Joseph K. Bradley authored 9 years ago

This modifies DecisionTreeMetadata construction to treat 1-category features as continuous, so that trees do not fail with such features. It is important for the pipelines API, where VectorIndexer can automatically categorize certain features as categorical.

As stated in the JIRA, this is a temp fix which we can improve upon later by automatically filtering out those features. That will take longer, though, since it will require careful indexing.

Targeted for 1.5 and master

CC: manishamde mengxr yanboliang

Author: Joseph K. Bradley <joseph@databricks.com>

Closes #8187 from jkbradley/tree-1cat.

7ecf0c46

[SPARK-9661] [MLLIB] minor clean-up of SPARK-9661 · a0e1abbd

Xiangrui Meng authored 9 years ago

Some minor clean-ups after SPARK-9661. See my inline comments. MechCoder jkbradley

Author: Xiangrui Meng <meng@databricks.com>

Closes #8190 from mengxr/SPARK-9661-fix.

a0e1abbd

[SPARK-9958] [SQL] Make HiveThriftServer2Listener thread-safe and update the... · c8677d73

zsxwing authored 9 years ago

[SPARK-9958] [SQL] Make HiveThriftServer2Listener thread-safe and update the tab name to "JDBC/ODBC Server"

This PR fixed the thread-safe issue of HiveThriftServer2Listener, and also changed the tab name to "JDBC/ODBC Server" since it's conflict with the new SQL tab.

<img width="1377" alt="thriftserver" src="https://cloud.githubusercontent.com/assets/1000778/9265707/c46f3f2c-4269-11e5-8d7e-888c9113ab4f.png">

Author: zsxwing <zsxwing@gmail.com>

Closes #8185 from zsxwing/SPARK-9958.

c8677d73

[MINOR] [SQL] Remove canEqual in Row · 7c7c7529

Liang-Chi Hsieh authored 9 years ago

As `InternalRow` does not extend `Row` now, I think we can remove it.

Author: Liang-Chi Hsieh <viirya@appier.com>

Closes #8170 from viirya/remove_canequal.

7c7c7529

Aug 13, 2015

[SPARK-9945] [SQL] pageSize should be calculated from executor.memory · bd35385d

Davies Liu authored 9 years ago

Currently, pageSize of TungstenSort is calculated from driver.memory, it should use executor.memory instead.

Also, in the worst case, the safeFactor could be 4 (because of rounding), increase it to 16.

cc rxin

Author: Davies Liu <davies@databricks.com>

Closes #8175 from davies/page_size.

bd35385d

[SPARK-9580] [SQL] Replace singletons in SQL tests · 8187b3ae

Andrew Or authored 9 years ago

A fundamental limitation of the existing SQL tests is that *there is simply no way to create your own `SparkContext`*. This is a serious limitation because the user may wish to use a different master or config. As a case in point, `BroadcastJoinSuite` is entirely commented out because there is no way to make it pass with the existing infrastructure.

This patch removes the singletons `TestSQLContext` and `TestData`, and instead introduces a `SharedSQLContext` that starts a context per suite. Unfortunately the singletons were so ingrained in the SQL tests that this patch necessarily needed to touch *all* the SQL test files.

<!-- Reviewable:start -->
[<img src="https://reviewable.io/review_button.png" height=40 alt="Review on Reviewable"/>](https://reviewable.io/reviews/apache/spark/8111)
<!-- Reviewable:end -->

Author: Andrew Or <andrew@databricks.com>

Closes #8111 from andrewor14/sql-tests-refactor.

8187b3ae

[SPARK-9943] [SQL] deserialized UnsafeHashedRelation should be serializable · c50f97da

Davies Liu authored 9 years ago

When the free memory in executor goes low, the cached broadcast objects need to serialized into disk, but currently the deserialized UnsafeHashedRelation can't be serialized , fail with NPE. This PR fixes that.

cc rxin

Author: Davies Liu <davies@databricks.com>

Closes #8174 from davies/serialize_hashed.

c50f97da

[SPARK-8976] [PYSPARK] fix open mode in python3 · 693949ba

Davies Liu authored 9 years ago

This bug only happen on Python 3 and Windows.

I tested this manually with python 3 and disable python daemon, no unit test yet.

Author: Davies Liu <davies@databricks.com>

Closes #8181 from davies/open_mode.

693949ba

[SPARK-9922] [ML] rename StringIndexerReverse to IndexToString · 6c5858bc

Xiangrui Meng authored 9 years ago

What `StringIndexerInverse` does is not strictly associated with `StringIndexer`, and the name is not clearly describing the transformation. Renaming to `IndexToString` might be better.

~~I also changed `invert` to `inverse` without arguments. `inputCol` and `outputCol` could be set after.~~
I also removed `invert`.

jkbradley holdenk

Author: Xiangrui Meng <meng@databricks.com>

Closes #8152 from mengxr/SPARK-9922.

6c5858bc